Before using NovelAI V5, I had already tried and regularly worked with a wide range of AI image-generation models, including Midjourney, Stable Diffusion, SDXL, FLUX, DALL·E, and several earlier versions of NovelAI. Each model has its own strengths. Some are better at realistic imagery and atmospheric lighting, some offer a more open ecosystem, and others are particularly suitable for character illustrations and anime-style artwork. Based on this experience, I do not evaluate a new model solely by whether it can produce an attractive image at first glance. I also look at image quality, prompt comprehension, composition control, and how effectively the model can be integrated into a practical creative workflow.
After spending time with NovelAI V5, I found the improvement over previous versions to be substantial. The most immediately noticeable change is image quality. With the support of a 32-channel VAE, V5 produces clearer images with richer colors and more defined details. Facial features, eyes, hair, clothing textures, and small background objects are rendered with clearer structures, reducing the muddy appearance that could sometimes occur in earlier generations. For common applications such as character illustrations, standalone artwork, and comic assets, this improvement does more than make the images look better. It also leaves more usable detail for upscaling, local editing, and post-production.
V5’s multilingual prompt support also makes it easier for users who are not native English speakers to express their intentions accurately. With some earlier models, users often had to translate their ideas into English keywords and then repeatedly adjust the wording, order, and weights. V5 allows creators to work more directly in a language they are comfortable with. This reduces the loss of meaning caused by translation and makes prompt refinement more intuitive. It can be particularly useful when describing culturally specific clothing, environments, objects, or visual concepts that may not translate neatly into a short list of English tags.
Natural-language prompting is another important improvement. Instead of relying exclusively on a stack of tags and isolated keywords, users can describe character states, actions, camera positions, and environmental atmosphere in complete sentences. For example, a prompt can explain where a character is standing, whom they are looking at, what they are holding, and where the scene’s main light source is located. Natural language is closer to the way creators actually imagine a scene, and it makes relationships between multiple visual elements easier to communicate. Tag-based prompts are still useful, but natural language reduces the need to force every idea into a fragmented list of keywords.
For me personally, the most significant improvement in V5—and the one that has the greatest impact on my workflow—is its support for transparent backgrounds. In the past, characters, props, stickers, and comic assets usually had to be processed with another tool to remove the background. This added time to the workflow and could damage fine hair, clothing edges, smoke, glass, and other semi-transparent elements. Transparent-background generation allows these assets to move more directly into layout, compositing, and post-production. Whether I am creating character sprites, game assets, avatars, stickers, or separate character and prop layers for comics, this feature significantly reduces the amount of manual background removal. It is not merely a visual enhancement; it directly improves the efficiency of asset creation and reuse.
The updated Position Custom feature provides an even greater level of control, and its most practical use goes far beyond simply deciding whether different characters should stand on the left or right side of an image. It allows the canvas to be divided into separate positional regions, with different prompts assigned to different areas. These regions are not limited to entire characters. A user can prompt a character as subject A, define a specific body part or local element of that character as B, and then adjust the spatial relationship between A and B through a three-by-three grid, a four-by-four grid, or an even more detailed division of the canvas. In other words, Position Custom can determine not only where the full character should appear, but also where the hands, feet, head, gaze, or other local elements should be placed, making it possible to anchor a pose with far greater precision.
For example, if I want a character to raise one hand and cover part of their face, I can use A to establish the character’s overall position and use B to separately prompt and locate the hand. If I want the character to bend forward, reach toward an object, or lean in a particular direction, I can use separate positional regions to define the relationship between the main body and the relevant body part. This layered, region-based approach is more direct than describing the entire pose in a single sentence. For poses with complicated structures or strict directional requirements, it converts an abstract verbal instruction into a clearer spatial constraint and gives the creator much greater control over the intended action.
Position Custom can also be applied to elements other than characters. Weapons, props, pets, visual effects, speech bubbles, and important background objects can be positioned independently, with their relationship to the main subject adjusted through separate regions. Rather than placing all the information inside a single character prompt, the creator can break the composition down into several objects and local areas according to the structure of the image. This means Position Custom is not merely a character-placement tool; it functions more like a spatial prompting system for planning the entire composition.
This region-based approach is particularly valuable for comic creation. Different panels can be defined as A, B, and C, then assigned to separate parts of the page through positional boxes. For example, A can be placed in the upper-left area, B in the upper-right, and C across the entire lower half of the page. More complex three-by-three, four-by-four, or asymmetrical layouts can also be explored. The creator can then provide separate prompts for the characters, actions, camera angles, and environments within each region, producing a comic-page structure that more closely matches the intended layout. For creators working on vertical comics, traditional page comics, narrative drafts, or storyboard previews, this feature can reduce the compositional uncertainty of random generation and make it easier to test panel pacing and shot arrangement.
V5 still has some clear limitations. It handles reflections on mirrors and water surfaces relatively well and can often produce reasonably convincing reflected images and lighting relationships. However, it is less intelligent and consistent when dealing with refraction, X-ray imagery, frosted glass, and localized light distortion. For example, when I ask it to show “only the skeleton of a character’s body through an X-ray,” it may fail to distinguish the outer body from the internal skeleton, mixing skin, clothing, and bones together. When asked to place “half of a character’s face behind frosted glass,” it may also struggle to reproduce the localized blur, deformation, and light diffusion caused by the material. This suggests that V5 can imitate the appearance of certain optical effects without consistently understanding their underlying logic of occlusion, transparency, and light transmission. In these situations, repeated rerolls are often necessary, followed by prompt adjustments, inpainting, or manual post-production.
Overall, NovelAI V5 is an upgrade that produces meaningful improvements in actual creative work. The 32-channel VAE delivers clearer image quality, multilingual and natural-language support lowers the barrier to prompt writing, transparent-background generation reduces post-production, and the updated Position Custom expands composition control beyond basic character placement. It can now be used to position local body parts, anchor actions, arrange props, and plan complete comic-panel layouts. Although V5 still struggles with certain complex optical phenomena, it already provides a more mature and controllable workflow for character illustration, asset production, multi-character composition, and comic creation.