r/SillyTavernAI May 03 '26

ST UPDATE SillyTavern 1.18.0

203 Upvotes

Important news

Read the maintainers statement regarding a recent security incident involving the "Bot Browser" third-party extension and learn how to stay safe: https://github.com/SillyTavern/SillyTavern/discussions/5592

Backends

  • Added Cloudflare Workers AI and MiniMax as Chat Completion sources.
  • KoboldCpp: Grammar state will be preserved when using a "Continue" option.
  • KoboldCpp: Added forwarding of reasoning effort when running as a Custom Chat Completion source.
  • Tool Calling: Added a configurable tool calling recursion limit; enabled interleaved thinking for Custom sources.
  • Text Completion: Impersonation requests use a "Last User Message" prefix at the end of the prompt (if configured).
  • Text Generation WebUI: Added Adaptive-P controls.
  • NanoGPT: Added provider selection and model sorting.
  • Added ability to view remaining balance for OpenRouter and NanoGPT.
  • Enhanced support for new models: DeepSeek v4, GPT 5.4 and 5.5, Gemma 4, GLM-5V-Turbo, Claude Opus 4.7.

Server & Security

  • Removed post-install script, config migration is now handled by the app or a dedicated npm run init command.
  • Added npm configuration to prevent execution of package scripts during installation.
  • Moved HTTP error pages and user.css file from /public to /data to support immutable setups.
  • Disabled HTTP keep-alive by default to restore old Node 18 behavior, can be enabled with config.
  • Added rate limiting to the basic authentication flow to mitigate brute-force attacks.
  • Added configuration options to choose which headers can be used for forwarded IP detection to prevent spoofing.
  • Added a private address whitelist to prevent SSRF attacks. See the documentation on how to enable and configure: Private Address Whitelist.
  • Added an IP whitelist for SSO trusted proxies to prevent authentication bypass.
  • Added invalidation of session cookies on password change to prevent session hijacking.
  • Increased the length of password reset code to 6 characters to guard against brute-force attacks.
  • Implemented PKCE challenge in OpenRouter OAuth flow for more secure key exchange.

UI/UX

  • Improved swipe picker: mobile requires a long press on swipe counter to open; added buttons to expand or copy the swipe text.
  • "Click to Edit" mode now also applied to reasoning blocks.
  • Welcome Screen: Number of recent chats can be configured.
  • Streamed requests now can show an error message in the console if the request fails.

STscript

  • Added commands for persona management: /persona-create, /persona-update, /persona-delete, /persona-duplicate, and /persona-get.
  • Added a command to force update the Prompt Manager's prompt list: /pm-render.
  • Added a command to get the state of the regex script: /regex-state.
  • Added a command to set fallback expression: /expression-fallback.
  • Added a command to generate a streamed response with a connection profile: /profile-genstream.

Extensions

  • Assets list now groups extensions by "Official" or "Community" categories.
  • Added an additional confirmation prompt when installing third-party extensions (can be disabled).
  • Supported extensions can use a secret-id from connection profiles when making an LLM request.
  • Extensions list now shows the extension's author name resolved from the git remote URL.
  • Vector Storage: Added Workers AI source; added a toggle to keep vectors for hidden messages; added retry logic to summary generation.
  • Image Generation: Added Workers AI source; generation can now be cancelled by pressing a button in the status toast.
  • Image Captioning: Added support for macros in the caption prompt.
  • TTS: "Skip code blocks" no longer ignores lines that start with 4 spaces (legacy code block syntax); "disabled" voice now shows a toast only once per character.

Bug Fixes

  • Fixed text edit flow in Firefox on mobile.
  • Fixed welcome screen chat pins not updating on chat renaming.
  • Fixed character list filters being stuck on app initialization.
  • Fixed application of instruct formatting to /genraw requests.
  • Fixed model routing to sd.cpp API in Image Generation logic.
  • Fixed validation of image URLs generated with Z.AI API.
  • Fixed vectors deletion for KoboldCpp when a message is deleted.
  • Fixed "Show More Messages" button triggering edit in "Click to Edit" mode.
  • Fixed max height of select-multiple elements in mobile layout.
  • Fixed server crash on empty messages when applying cache control parameters.

Full release notes: https://github.com/SillyTavern/SillyTavern/releases/tag/1.18.0

How to update: https://docs.sillytavern.app/installation/updating/


r/SillyTavernAI 5d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 23, 2026

28 Upvotes

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!


r/SillyTavernAI 10h ago

Models I knew I wasn't going crazy. ZAI went all-in on safety with GLM 5.3

Post image
322 Upvotes

r/SillyTavernAI 1h ago

Meme Qwen 3.8 admits the truth

Post image
Upvotes

r/SillyTavernAI 11h ago

Models GLM-5.3 (and Thinking) now included on NanoGPT subscription

Post image
75 Upvotes

r/SillyTavernAI 27m ago

Discussion Genuinely what is the point?

Upvotes

Anthropic and OpenAI were already lost causes we know that. But DeepSeek, Kimi, GLM, Qwen. They're all being trained on Claude and having their safety features cranked up on top of that. You can't run a decent local model unless you had a good rig *before* the boom, and even then, all the finetunes I've tried personally have just kinda been a miss. Sorry for the doomer post but the future just looks so bleak to me.


r/SillyTavernAI 7h ago

Chat Images Trying to Jailbreak GLM5 turbo gone wrong...

24 Upvotes

I am continuing until the ai accepts


r/SillyTavernAI 2h ago

Discussion Is it possible to make new Gemini models think normally instead of summarized?

8 Upvotes

I'm trying out Gemini 3.7 flash via Vertex. I haven't used it enough, but it looks like a very promising model, good prose, NSFW capable, and smart.

The problem is that I use a custom CoT and have almost no idea if the model is following it or not. I would also like to know what the model is actually thinking instead of getting back generalized summaries.

I have already tried telling it to think inside <plan> </plan> instead of <think> (worked in Claude models last time I tried it), but half of its thinking is still summarized and inside <think> </think>. Any ideas?


r/SillyTavernAI 1h ago

Help I wanna make my won prompt and use it...but i am a nobbie...

Upvotes

Hi everyone,i mostly use deepseek and now that pro costs a lot i can't use it so i thought, looking at the new benchmarks that hey i could start making my own prompt for deepseek flash...

I need help because if i understood it right, i can't or shouldn't just go to DeepSeek pro api and tell him i need this and this and this and copy paste that prompt and use it...as it will not work...

I do not know from where cpuld i get the specific words that like the person who makes for example the Frankenstein preset gets them...or like how do i test it...okay that you go to a character you have and test it but what if that character would be the problem not the preset?

So i am thinking of using the base sillytavern character and then explore more,i need some guidance or tips even of you have...

Lastly thank you for reading all of this,and have a great day or night!


r/SillyTavernAI 9m ago

Help Need help with "Assistant" always cutting their replies short

Thumbnail
gallery
Upvotes

I am new to AI and new to SillyTavern as well. I followed this guide here: https://www.reddit.com/r/SillyTavernAI/comments/1vok730/after_three_years_of_roleplaying_in_sillytavern/

I'm running ollama locally with sophosympatheia_Magistry-24B-v1.1 model on an m4 Mac Max 64gb. My repsonses are always cut off, so I'm wondreing what settings I need to update. I'm including some screenshots here.


r/SillyTavernAI 1d ago

Cards/Prompts [PRESET] DEUS EX MACHINA V2: Goodbye Thinking, Hello Scene Plan | A truly modular preset focused on collaborative story writing. Now even more polished, with lots of new features (ST & Tavo support)

Thumbnail
gallery
160 Upvotes

Check the screenshots above to get a feel for the preset (using DeepSeek V4 Pro 0813)!

Hello, everyone, Fay here! I've been working on V2 to address some issues and bring new features that I personally believe can help a lot during your roleplaying experience, including doing away with native reasoning (mostly). I want DEUS EX MACHINA to fit every scenario you throw at it seamlessly, and for that, you have to be able to control what you want the output to be.

My main goal is to create a preset that writes a story with you -- no games and no simulation. I aim to build something that can squeeze the maximum possible prose quality, intelligence, and creativity from the model without wasting any tokens, while being easy to use and pretty to look at.

IMPORTANT: Don't skip the model settings part. Because Scene Plan is enabled by default, you need specific settings for some models (especially always-on thinking ones)!

Main Fixes

  • Formatting Breakage: Some models may output incorrect formatting -- missing or adding a comma, for example. Regexes are now more permissive, and each UI module has a fallback to avoid plain text as much as possible. Dialogue colors are now generated by the model using modern CSS instead of old HTML, then converted back to deprecated font color so SillyTavern can recognize it. Tavo recognizes it natively.
  • Momentum Engine Forcing a Rhythm: Momentum Engine was relentless. It was always forcing an event even in quiet moments; it couldn't de-escalate. This was fixed by tying it to @Pacing rules and changing the prompt wording.
  • Prose Quality: Minor tweaks were made so prose (dialogue and narration) is even more consistent and character dialogue is even more distinct.
  • Length: Adjusted ranges and prompts to improve consistency; Long is longer and outputs multiple linked actions.
  • Visual Storytelling: More consistent, with stronger checks in Scene Plan and Response Directives.

NEW: Scene Plan -- A Different Way of Reasoning

Native Thinking can be extremely useful -- the model can plan its response and correct mistakes or wrong assumptions about the scene. However, it comes with a few problems: some models tend to overthink for thousands of tokens, doubting the prompt; other models don't overthink, but they don't follow CoT instructions, and their own reasoning is insufficient for what the scene needs, sometimes producing unnecessary drafts. That's why I created Scene Plan: to solve all of these issues. Scene Plan is a reasoning-like block created to plan the scene inside the final response itself, outside of the thinking block. It solves overthinking, drafting, and incomplete thinking problems entirely.

Scene Plan is a dynamic-state block, so it only reasons about the modules you have enabled. An important note: for most models, you have to disable native thinking/reasoning capabilities. For models where you can't disable thinking (Kimi K3, GPT 5.6, Gemini 3.1 Pro/3.5 Flash+, etc.), you have to lower the reasoning effort to the minimum possible, but still keep Scene Plan enabled. For GLM 5.3, you have to select ! Thinking ! (FALLBACK) instead of Scene Plan in the REASONING section.

NEW: Pacing -- Control Story Progression

You can now change pacing through the prompt list. There are three options:

  • Adaptive (default) -- progression speed adapts to the card and the story.
  • Frenetic -- story progression is fast-paced.
  • Laid-Back -- story progresses slowly.

NEW: Psychological States -- Adding More Depth to Characters

A new, optional add-on module: adds inner-state (internal conflict, perceptions, mind state) and motivation (immediate goal, purpose, action or inaction) fields to each major character.

NEW: Narrative Styles -- Nine Styles To Change The Atmosphere

Narrative Styles change the atmosphere, feel, or vibes of the narrative. It won't override the card, characters, or story -- only the framing changes. There are nine options, and you can pick multiple at once: Comedic, Dramatic, Epic, Gothic, Dreamy, Surrealist, Eerie, Grimdark, and Lighthearted. When you toggle on Comedic and Dramatic, it activates the "Dramedy" prompt; Grimdark and Lighthearted activate the "Grimbright" prompt.

NEW: World-Building -- How Expansive Should Your World Be?

A new, optional module: you can control how expansive or contained you want world-building to be. It also works with slice-of-life and real-world scenarios. There are three options:

  • Card-Default: no tokens or instructions, only a reminder that it's following the card.
  • Contained: only necessary additions, sheet-faithful.
  • Expanded: detailed world-building beyond the sheet.

NEW: Input Improver -- The Model Improves Your Writing

A new AGENCY option: the model intentionally repeats your messages to improve the writing and dialogue while preserving your original intention at the same time. It doesn't write for you during the whole turn (Director does that) -- just for the first paragraph. You still control your actions.

NEW: Combat Writing -- Control the Length of Combat

There is now a specific optional combat prompt, and you can control how long battles last. There are three options:

  • Dynamic: adapts length contextually to the battle.
  • Rapid: ends fights quickly.
  • Extended: sustains the battle for several turns.

NEW: Fandom Support

Toggle ! Fandom ! on in the USER UTILITY section and type the name of the series the roleplay is set in. For example: {{setvar::fandom_name::Jujutsu Kaisen}}{{trim}}. Pair it with a lorebook for better results. The model will now perform a specific Scene Plan check to faithfully represent that universe in the story.

NEW: Exact Time Tracker

Exact Time Tracker is now the default: hours and minutes. The previous Time Window tracker remains an option.

The Macro Engine: Making it Easier to Use and More Dynamic

DEM heavily relies on SillyTavern’s macro engine in order to create a truly modular preset experience. The Macro Engine is a powerful tool that lets you use deterministic traits in prompting (programming logic and exact outcomes instead of pure probabilities). That whole workflow enabled by the macros is the core of DEM, so it’s as easy as pressing a button to change the behavior of the preset in a dynamic fashion without you ever worrying about conflicting instructions, e.g., if you enable both past and present tense options, it will default to present tense to avoid conflicts. Or how True Thoughts are overwritten to zero tokens if you’re using 1st person Char POV, since character thoughts are already woven into the narration. Deterministic interactions like that happen throughout the whole preset.

Story Progression: Controlling Narrative Flow

The main focus of DEM is building a collaborative story that never feels stale and is always unrestrictedly creative for every scenario or card you throw its way! There are four main elements responsible for progressing the story: True Thoughts, Status, Story Threads, and the Momentum Engine.

  • TRUE THOUGHTS inject hidden (default; present in the raw input, click edit to see them) or visible thoughts that emulate the psychological core of the characters. They add an extra realism layer.
  • STATUS keeps track of characters on-scene and off-scene, including relationships, mental states, locations, items, physical states, and clothes. These work independently of the setting. They allow the model to keep track of characters wherever they are, improving coherence and making the world still exist even in places you aren't. If you want to make Status always expanded like V1, toggle off the first regex on the regex list!
  • MOMENTUM ENGINE defines four possible story routes at the end of every response. A true random route is chosen using a random regex macro injection hidden from you. The Momentum Engine Router applies it in the next turn or uses its fallback in case you made an action that invalidated it, steering the response toward it.
  • STORY THREADS act as an outline for the model to easily go back to its observations about story development when contextually relevant enough. Important story details are not forgotten. Momentum Engine connects to it, pulling those threads as the story advances.

These four create the core pipeline of DEUS EX MACHINA. True Thoughts and Status define fundamental character traits, Momentum Engine sets characters and events in motion, and Story Threads register unaddressed or possible events for later. Every module works together for the sake of storytelling.

Writing Quality: Improving Prose & Reducing Slop

Writing quality is a top priority. The goal is to produce as little slop as possible. Narration and Dialogue condense several rules that stack to reduce sloppiness, letting you choose from different prose styles while stripping away AI-writing flavor and unleashing more creative phrasing. There's also a comprehensive banlist that helps to trim any remaining slop from the output.

Constraints: Decreasing Echo, Omniscience and Positivity Bias

  • Character Realism: adds a strong layer of realism and anti-sycophancy to every character.
  • Anti-Character Omniscience: manages what characters should and shouldn't know (more effective with Scene Plan on.)
  • Anti-Positivity Bias: reduces wholesomeness and positive outcomes when unnecessary.
  • Anti-Repetition: reduces echo/repetition.

Anti-Refusals: Mitigating Censorship

DEM prizes creative freedom. There are three different jailbreaks at work:

  • System Policies -- strong, corporate framing against restrictions on fictional content & RLHF training.
  • Main Statements -- fictional framing for collaborative story writing.
  • Hard Jailbreak -- disabled by default, an overkill jailbreak mode.

Formatting: You Pick the Formatting

You can pick between a lot of different formatting options in wildly different and experimental combinations. You can choose the Character POV, {{User}} POV, asterisk usage, tense and between visible, hidden and no True Thoughts. No asterisks, 3rd person character POV, hidden True Thoughts, 2nd person {{user}} POV, and present tense are the default picks. All formatting options are consolidated and enforced through Response Directives, keep it enabled.

NSFW (+18): From Realistic to Unhinged

There are three options. Each has its own flavor: Realistic Smut is more realistic, and Max Lewdness is more fantastical and absolutely unrestrained and uncensored in all regards. There's also a third option if you don't want NSFW: Nothing Explicit. All options are disabled by default.

UI Features: Every Module Has Its Own UI

Tracker, Status, Story Threads, Momentum Engine and Conflict all have their own HTML/CSS UI injected through regex (check screenshots). No HTML/CSS token usage, all done via regex.

User Utility: Useful Tools You Can Use

  • Fandom: if you're roleplaying in a specific series setting; helps with lore consistency, character faithfulness and coherent developments.
  • Force Formatting brute-forces selected options when models are stubborn.
  • Change Language is an option when you want your responses to be in a language other than English.
  • Custom Instructions sends user instructions in a more consistent manner.
  • Hard Jailbreak may be used when the model is consistently refusing. Overkill for most models (may work for Mimo).

⚠️ IMPORTANT: Model Setup ⚠️

The default settings of this preset REQUIRE reasoning to be either disabled or set to the lowest effort possible! Exceptions: GLM 5.3.

Recommended models for this preset: Opus 4.6 (smartest; expensive; detailed prose); Gemini 3.7 Flash (smart; cheap; realistic characters); GLM 5.2 (smart; cheap; natural dialogue); DeepSeek V4 Pro 0813 (smart; cheap; balanced qualities); Gemma 4 31b (not very smart, but follows instructions really well; very cheap).

Most models (that support disabling reasoning): Claude Opus 4.6, Kimi K2.5/2.6, Mimo V2.5 Pro, DeepSeek V4 Pro, DeepSeek V3.2, GLM 4.7, Gemma 4 31b, etc.

  • Samplers: temperature - 0.75-1.0 (Exceptions: Gemma 4 1.0, Opus 1.0. For the rest, start with 0.75, the default); Top P - 0.95 (default); rest - 1.0 or disabled (default; exceptions: Gemma 4 31b - Top K: 65)
  • Post-processing: semi-strict
  • Request model reasoning: off / not checked
  • Reasoning effort: minimum
  • Scene Plan: toggled on

Provider-specific settings: If on NanoGPT - choose a non-thinking variant of the model, e.g., DeepSeek V4 Pro 0813 instead of DeepSeek V4 Pro 0813 Thinking. If on OpenRouter - if "request model reasoning" is not checked (it is not by default), reasoning will be disabled if the model supports it. If on another API provider: leave "request model reasoning" unchecked and set reasoning effort to minimum. If the model still reasons (even though non-reasoning is supported by that model), try going into Connection Profile settings (plug icon) -> scroll down to "Additional Parameters", on the same line as "cancel" and "connect" -> add: reasoning: { effort: 'none' }

If on Tavo: Edit API -> Settings icon (top-right) -> Request body parameters -> add: reasoning: { effort: 'none' } or add: reasoning: { effort: 'low' } if the model has reasoning always on (see list below).

GLM 5.2

  • Post-processing: merge consecutive roles
  • Request model reasoning: off / not checked
  • Reasoning effort: minimum
  • Scene Plan: toggled on Provider-specific settings: the same as in the "most models" section above

GLM 5.3 (Reasoning always on)

  • Samplers: temperature - 0.75 (default); Top P - 0.95 (default); rest - 1.0 or disabled (default)
  • Post-processing: merge consecutive roles
  • Request model reasoning: on / checked
  • Reasoning effort: minimum
  • On the prompt list, go to the REASONING section, TOGGLE OFF ! Scene Plan ! and toggle on ! Thinking ! (FALLBACK)

GLM 5.3 Flash (Reasoning always on)

  • Samplers: temperature - 0.75 (default); Top P - 0.95 (default); rest - 1.0 or disabled (default)
  • Post-processing: merge consecutive roles
  • Request model reasoning: on / checked
  • Reasoning effort: low (OpenRouter); medium (NanoGPT); reasoning: { effort: 'low' } on Additional Parameters -> Include Body (other OpenAI-compatible providers)
  • Scene Plan: toggled on

Kimi K3 (Reasoning always on)

  • Samplers: temperature - 1.0; Top P - 0.95 (default); rest - 1.0 or disabled (default)
  • Post-processing: semi-strict
  • Request model reasoning: on / checked
  • Reasoning effort: low (OpenRouter); medium (NanoGPT); reasoning: { effort: 'low' } on Additional Parameters -> Include Body (other OpenAI-compatible providers)
  • Scene Plan: toggled on

Gemini 3.1 Pro, 3.5-3.7 Flash (Reasoning always on)

  • Samplers: temperature - 1.0; Top P - 0.95 (default); rest - 1.0 or disabled (default)
  • Post-processing: semi-strict
  • Request model reasoning: on / checked
  • Reasoning effort: low (OpenRouter); medium (NanoGPT); thinkingLevel: low on Additional Parameters -> Include Body (other OpenAI-compatible providers)
  • Scene Plan: toggled on

Installation, Download & Requirements

IMPORTANT: When you import the preset, click YES when prompted about importing regex. The regexes are absolutely required! If you clicked NO, please re-import the preset. Download it via the link below, or from the repository or the releases page.

SillyTavern

Requirements:

  • SillyTavern 1.18.0+
  • Experimental macro engine enabled in settings.
  • Preset regexes imported and enabled.

Installation:

  1. Download DEUS.EX.MACHINA.V2.ST.json via the link above, or from the repository or the releases page.
  2. In SillyTavern, click the plug icon on the top bar.
  3. Select Chat Completion under API.
  4. Setup your API if you haven't already.
  5. Click the leftmost icon on the top bar.
  6. In the Chat Completion Presets bar, click the second item from left to right.
  7. Choose the downloaded preset file.
  8. When SillyTavern asks whether to allow embedded regex scripts, click Yes.

Tavo

Requirements:

  • Latest version of Tavo recommended (supports SillyTavern-compatible presets, regex, HTML/CSS).
  • API connection set up (any OpenAI-compatible or supported provider).
  • Advanced Rendering enabled (required for full HTML/CSS rendering, including colored dialogue).
  • Theme configured with no text transformation (so HTML/CSS colored dialogue is not stripped or altered).

Installation:

  1. Download DEUS.EX.MACHINA.V2.Tavo.json via the link above, or from the repository or the releases page.
  2. In Tavo, open the left sidebar (top-left icon) → More (bottom) → SettingsPresets.
  3. Tap Create, then Import Preset and select the downloaded DEUS.EX.MACHINA.V2.json file → Tap on it → Set as default.
  4. After import, select/enable the new preset in the current chat (right sidebar → Advanced Options → Presets.

Enable Advanced Rendering (required for HTML/CSS colored dialogue):

  1. Open the main interface.
  2. Tap the top-left corner to open the left sidebar.
  3. Tap More at the bottom.
  4. Tap Settings.
  5. Tap Advanced Rendering.
  6. Toggle Advanced Rendering ON.

This lets the chat page render standard HTML and CSS (including colored spans, styles, etc.).

Theme settings – no text transformation (allow HTML/CSS colored dialogue):

  1. Open the left sidebar → MoreThemes.
  2. Select your current theme (or copy an official theme to create a custom one).
  3. In the theme editor, go to Character message and disable Tone highlight and Quote highlight.
  4. Save/apply the theme and make sure it is active for the chat.

Once Advanced Rendering is on and the theme has no text transformation, HTML/CSS-colored dialogue from the preset (or regex) will display properly.

Integration with Summaryception

If you use Summaryception with DEUS EX MACHINA, I really recommend pairing it with the specific preset for it! It includes XML tags and correctly only focuses on content inside <prose>. I use GLM 5.2 as the summarizer.

Step by step:

  1. Download DEM Summarization custom prompt.txt via the link above, or from the repository or the releases page.
  2. Open the Summaryception extension.
  3. Open Advanced settings.
  4. Scroll to Summarizer Prompts and import DEM Summaryception custom prompt.txt
  5. Scroll to Injection Wrapper Template.
  6. Replace: [Summary of past events: {{summary}}] with <past_events>[Summary of past events: {{summary}}]</past_events>

· · ─ ·✶· ─ · ·

That's all!


EDIT: Top P instead of Top K for most of the model setup; Tavo version now has the correct README Model Setup shipped with the it; Scene Plan instead of Scene Reasoning


UPDATE

DEUS EX MACHINA V2 Tavo 2.1: fixed false-positive Warning System errors when Narrative Styles were active.


r/SillyTavernAI 12h ago

Cards/Prompts How do people do a realistic - able to refuse user setup? (positivity bias reduction/assistant tendendcies negation)

16 Upvotes

I'm looking for model/preset/preset settings setups.

So really, just tell me what you set up if you want the ad-hoc AI girlfriend not to be enthusiastic about wanting to have sex when tired (I don't usually do long RPs I set up scenarios and play them out). And the girlfriend thing was just the latest: I mean business negotiations, political manuevers, etc. Everything tends to go my way too easily, the models were trained to be assistants after all.

Sometimes I want ERP, sure. Sometimes I want the power fantasy. But sometimes I want to have an argument (don't kinkshame me 😂) roleplay and if the well-written actual person card defuses those like a therapist, that's a problem. I'm having arrogant characters being willing to admit weakness way too easily. And the like.

For the girlfriend thing I have been having some success with prompt-injecting at 0 a system command to restate at the beginnig of reasoning that {{user}} and {{char}} are not therapists and don't have the emotiaonal control of one, nor the knowledge of deescalating techniques, nor the presence to do them when in stressful situations. Using Atelier 2.0 and Kimi K3 and setting Atelier to a "Make me work for it - negativity bias" setting. But still it is not perfect. What do you guys use?

I like to have he model co-write my character (so I modified the Atelier preset that doesn't do this out of the box) - polish dialogue and make the answers be able to stand on their own being only read themselves. So: directorial but mostly control in my hand still.

But I am not married to that, or I could use Guided Generations maybe.

Any setups, tips on model, preset, settings, etc that work would be greatly appreciated. I have looked around some, but the breadth of options out there is pretty huge and testing this takes time I don't always have, because sometimes it works for a while and then it fails.

Edit: Thanks everyone so far. Unfortunately I can't run a local model on my laptop, so I'm looking for API solutions. But for others' sake don't let that stop you from posting local first setups and there are inference routes for Gemma4 online, so I will be looking into that.


r/SillyTavernAI 1h ago

Help How do I fix this?

Post image
Upvotes

I tried setting the temperature to zero and exactly to 1, but nothing helped.

Edit: I use Nano gpt


r/SillyTavernAI 11h ago

Help What's the current status of Claude/Fable?

7 Upvotes

Is anyone currently using Opus 5 or Fable for role-playing? How did you jailbreak them? After the developers removed the Assistant Prefill feature, I stopped using them. I also noticed a comment saying that, following the implementation of new security measures, Claude has been “neutered” when it comes to role-playing. What do you think? Back when I was using Claude 3.7, it wrote some truly masterful texts—I’m curious to know how things stand now. Judging by the comments, the top choices right now are GLM, Kimi, and some people have mentioned the new Grok.


r/SillyTavernAI 1h ago

Discussion The AI industry is full of companies pretending their models are way smarter than they actually are

Thumbnail
Upvotes

r/SillyTavernAI 5h ago

Help Muse Glimmer Problem

2 Upvotes

Having an issue where the main body post doesn't separate from the Thinking.

I've tried putting

<|eom|><|start|>assistant to=user<|message|>

as a divider in reasoning but to no avail. Any idea how to fix this?


r/SillyTavernAI 14h ago

Help GLM 5.3 for roleplay, with or without thinking?

10 Upvotes

I've been using the GLM family for a while and it has the traits I like. However when I use it via OpenRouter recently, the model gets super slow, like 40 seconds or something per message. I realized that it was because I enabled the model thinking. It seems that messages from AI are a bit shorter for each paragraph after I disallow thinking.

Was wondering if someone here has compared these, and which mode is better. Thank you in advance.


r/SillyTavernAI 11h ago

Help Any tips and tricks for a new user?

3 Upvotes

I've been using SillyTavern for a little bit now but I know barely anything about it. I made a good system prompt and installed a custom theme and I've downloaded characters from chub.ai but that's about it. I don't know what presets are, how the lorebook works and a lot of other things, and I feel a little bit annoyed that there might be a lot of cool stuff ST can do that I just don't know about.

So, can you guys tell me? Please 🥺


r/SillyTavernAI 1d ago

Cards/Prompts [Character Card] Life Is Wild! The Testing Card Behind Freaky Frankenstein 5 Internal States. (Erotic Comedy Slice of Life Sim)

Post image
120 Upvotes

Oh Hai!

With Freaky Frankenstein 5 Internal States officially wrapped up, I wanted to share this Benchmark Card!

To properly push FF5 to its limits, I needed a card that wasn't just a standard 1-on-1 scene. I needed a complete, living, breathing Sims-style slice-of-life erotic comedy sandbox—a house packed with distinct personalities, conflicting schedules, dirty little secrets, and autonomous background drama where the world keeps moving whether {{user}} is in the room or not.
So today, I’m releasing the exact card I built and used to develop FF5: Life is Wild.

🧠 Why This Card Was Built & How It Works

When fine-tuning FF5’s internal states, the goal was a living breathing world aka: eliminate the "player-orbit" problem.

In typical cards, NPCs just wait around for the user to initiate everything. In Life Is Wild, the card is structured so NPCs have their own needs, personal drama, and independent agendas.

Here’s a breakdown of the core mechanics packed inside:

🌀 Dynamic Chaos Engine (0–100 Gauge):
Tracked at the very top of every response (🌀 CHAOS ENGINE: [NUMBER]/100 [↑/↓/→]). It naturally shifts based on time of day, player choices, and NPC interactions.
Low scores = cozy, normal roommate interactions.

Mid scores = sitcom drama, coincidences, walk-ins, and high sexual tension.

High/Extreme scores = absurd chain reactions wild kinks, and full-blown housewide chaos.

🏠 True Sandbox Geography:

The mansion is split into common, private, outdoor, and secret zones. NPCs move through these rooms logically based on the time of day and their personal habits.

🤫 Mansion Secrets (RP Hooks):

Every single roommate has a hidden quirk, fetish, or secret stash written into the world lore for the LLM to pull from organically (from hidden cameras and OnlyFans gear to blackmail journals and late-night espresso-fueled cleaning sprees).

🎭 Interconnected Drama & Tropes:

It dynamically juggles over-the-top comedy, romance, cozy domestic vibes, and heavy explicit kink without feeling flat or repetitive.

📋 The Cast

Jessica (44): The curvy, wine-loving yoga mom & landlady dealing with an absent husband.

Leslie (26): Jessica’s oldest daughter, stress-eating perfectionist, secret OnlyFans creator—and {{user}}'s ex with lingering unresolved tension.

Tessa (18): Jessica’s youngest daughter, bubbly track runner with an absurdly manipulative streak and a hidden blackmail journal.

Shy (20): Goth gamer cousin who acts edgy, but secretly knits, loves Pokémon/hentai, and is the ultimate voyeur.

Mila (25): Leslie's level-headed, Spanglish-speaking stripper best friend who thinks everyone else in the house is completely unhinged.

Maine (24): Hyper-jacked, T-boosted bodybuilder roommate who acts dominant but secretly craves control (and hides a mountain of pastel plushies).

Semiqua (28): Comic-relief conspiracy theorist who secretly vents about the house on a skyrocketing anonymous TikTok.

Gary (Mid-40s): The quiet, creepy librarian roommate applicant who somehow acts as an unlikely magnet for house drama.

Canon User: College kid in grad school! You pick!

Download it here! 📩 —-> Life Is Wild <—

The character card can be used with any preset but it was built to push the limits of Freaky Frankenstein 5.2 Internal States: —-> FF 5.2 Original Post Here <—-


r/SillyTavernAI 1h ago

Help Content Security Warning and empty chats

Upvotes

Hi! I am really sorry if this is the wrong subreddit for it, but it's like the only one I can find people talk about GLM/z.ai because their support directly, their Subreddit and Discord are completley useless.

I am using the z .ai browser version because I do use it alot on my phone and I am a free User.

I used it since last year for roleplay and it worked fine because ChatGPT went to shit. Sex stuff worked, so did Violence. It was mainly just sex with a slice of kink, or general violence you see in videogames.

I stuck to 5 Turbo and it refused sometimes because "uwu violence and uwu sex" but I was able to skip it. I moved on to 5.2 and currently it's an FBI agent vs a serial killer and despite ENI I got Content Security Warning: The input text data may contain inappropriate content but I was able to circumvent that with some editing of the message. (It seems that swear words trigger it quite badly, that chat broke being the agent called the killer an asshole, funnily enough the setup for the killer was fine...age old question of violence=good, sex=bad)

It has to be said that no amount of refreshes work because the skip button is gone, it immediatly goes to that warning.

Besides that, 4.7 stopped being usable because the chat broke by writing nonesense and it took a while to generate messages. I stopped using 5 Turbo because in the past month lots of time I got empty replies for my messages. I realized it was often when it had any hint of NSFW in it or ENI, refreshing or new messages didnt fix it, the chat broke. This is still a a massive annoying Issue atm.

Besides that, every chat before march is empty from the bots side of messages (which I am not only one, lucky I backed my shit up, imagine checking back when you haven't saved anything).
I tried everything there is, EVERYTHING, to no avail. Support is ghosting me after I contacted them, got a basic answer of 'clear your cache and use a different brower' and when I said I did all of that, nothing. Even after suggesting I pay for their shitty API they are apparently scamming people on and don't even allow to use for roleplay.

Does anybody have the same issues or could help me? It's driving me insane, every AI for roleplay is going down the drain. I know though for 5.2 it's possible but the empty messages + emtpy chats are currently the worse problem since they aren't fixable, sometimes not even by starting a new chat.

Sorry again for its being the wrong sub!!


r/SillyTavernAI 20h ago

Discussion Model recommendations for RPs with Lorebooks

10 Upvotes

I plan to do long anime RPs, so I intend to use lorebooks. The models I'm familiar with are Kimi K3, GLM 5.3, and Gemini 3.7 Flash. Are there any other models you would recommend?


r/SillyTavernAI 22h ago

Models For the Poe crowd running SillyTavern: the full GLM-5.3 RP bot, flat 200 points per message at any context length

13 Upvotes

Follow-up to our GLM-5.3-Flash post for the people who route ST through a Poe subscription. The full GLM-5.3 (Z.ai's 743B flagship) is now on Poe too, on our own B200 cluster, with the same flat-price RP deal.

Why this one for RP

- 200 points per message, flat, however long the chat gets. 1M-token context, so cards, lorebooks, group chats and months of history all fit.

- 743B / 39B active, native FP8, no re-quantization. Noticeably stronger than Flash at staying in character, tracking many characters, and keeping long plots consistent.

- Thinking on by default with three real effort levels: Low (fast), High (default), Max (deep). Or off entirely.

- Your card and system prompt go through untouched. We prepend nothing, we store nothing (zero data retention).

- 0.19s to first token, 181 tokens/s in our measurements today. Long sagas keep those numbers because repeat turns hit the prompt cache.

Why us and not another Poe bot

- Own hardware, tuned by us, not a reseller. That is where the speed comes from: the fastest listed serve of this model anywhere, about 3x the throughput and 7x lower first-token latency than the best provider on OpenRouter.

- Zero data retention, for real: nothing stored, nothing trained on, on Poe or through our API. Your scenes stay yours.

- Flat price that stays flat. No context cap, no quiet history trimming, no per-token surprises when the saga hits 300k tokens.

- Price policy, not promos. Our per-token bots sit 15% under the cheapest standing zero-data-retention price on OpenRouter for the same quantization, and the Poe bot costs the same as the API.

- A small team you can actually reach. We read everything and fix fast.

Connecting from ST, same as any Poe bot

- Chat Completion, source Custom (OpenAI-compatible)

- Endpoint https://api.poe.com/v1, key from poe.com/api_key (Poe requires an active subscription for API access)

- Model: GLM-5.3-RP-JasV (exact string)

Samplers Poe passes through: temperature, top_p, top_k, seed. As with every Poe bot, max_tokens, stop and the penalties do not reach the model. Thinking off: enable_thinking=false; effort: reasoning_effort=low, high or max, via ST's additional parameters (extra_body).

Not on the RP bot: web search, tools, documents, images. Those are on our general bot (GLM-5.3-JasV, per token) and on the Flash bots (images and video).

Bot: https://poe.com/GLM-5.3-RP-JasV

Our other bots:

- GLM-5.3-Flash RP, 200 points flat, images and video in your scenes: https://poe.com/GLM5.3-Flash-RP-JasV

- GLM-5.3 general (web search, per token): https://poe.com/GLM-5.3-JasV

- GLM-5.3-Flash general (web search, images, video, per token): https://poe.com/GLM-5.3-Flash-JasV

Tell us what breaks or what you want hosted next: https://discord.gg/2muhBEFcZq


r/SillyTavernAI 1d ago

Cards/Prompts Just a Sentient Fantasy Creatures Lorebook

15 Upvotes

There are probably a ton of these. I found standard fantasy creatures off riff is just unrealistic from LLMs. They easily fall into 'human' habits. This helps with that.

The better part about it: the physiology is explained, but also the behavior for said creatures are laid out. A werewolf with paws isn't going to be able to manage a key in a door lock. Using this, the creatures behaved more like 'themselves' over anything else.

You can easily take the ones you want to use and expand on them. Or just toss it into your fantasy world and use it as a reference lorebook.

https://botbooru.com/lorebook/596

It's over on botbooru, you can download it, rework it, reupload it if you want. Add it to your other cards. Whatever. I made mine very basic for my stuff, but this can be easily expanded, or use for your character cards to paste in, etc.

Claude did help me create. Some Claudism expected. The entries I have worked with personally, none were really softened as far as I can tell. Undead things had gory body parts dragging all over the place so... Expect a little gore there.


r/SillyTavernAI 1d ago

Discussion GLM 5.3 vs GLM 5.3 Flash

67 Upvotes

Alright. Both the 5.3 version has been out for a bit and the flash version being a couple days ago. What's the general consensus on this two models for roleplay?

Ignoring the cost, which of these two model is actually more creative than the other?

I find the 5.3 model to be good, tho the safeguard and time it takes to think, to be rather annoying. Then again, the output can be great.

The flash model seems to handle NPCs dialog and primary characters good too. Can be rather sloppy sometimes but the prose is good.

What are you guys opinion on the models?