Illustrious XL

Illustrious XL - Open source SDXL model for native 1536px anime art generation

Launched today

Standard AI image models struggle with native high-resolution anime art, often requiring external upscalers that degrade quality. Illustrious XL is an open-source SDXL-based model that natively generates 1536x1536 images without post-processing. Its hybrid prompting system combines natural language with Danbooru structured tags for precise creative control. The untuned base canvas serves as an ideal foundation for LoRA training and custom checkpoint development. Built by OnomaAI Research under permissive open-source licenses, it powers thousands of community-derived models.

AI ImageFreemiumNLPImage GenerationStable DiffusionCustom TrainingOpen Source

What Is Illustrious XL

Anime AI creators face a persistent trilemma when working with standard Stable Diffusion XL (SDXL) models: output resolution caps at 1024px, forcing reliance on Hires.fix or external upscalers that introduce blurry lines and distorted details; natural language prompts fall short of precisely describing anime-specific attributes like hairstyles, eye colors, and costume elements; and most base models come heavily aesthetic-tuned, making them unsuitable as clean canvases for LoRA training or custom checkpoint fine-tuning.

Illustrious XL (ILXL) is an open-source AI image generation model built on the SDXL architecture, developed by OnomaAI Research based in South Korea. It was purpose-built to solve these exact bottlenecks. The model directly addresses three fundamental technical limitations:

  • Native 1536×1536 resolution generation — trained at ultra-high resolution, not post-processed upscaling
  • Hybrid Prompting system — seamlessly fuses natural language (NLP) with Danbooru-structured tags
  • Untuned Base Canvas — deliberately left free of aesthetic bias, offering the cleanest foundation for LoRA and checkpoint training

These three differentiators position Illustrious XL as what the community now calls "a new cornerstone of open-source AI illustration." On platforms like Civitai, thousands of derivative models and compatible LoRAs trace their lineage back to ILXL checkpoints. The model has undergone aggressive iterative evolution — from v0.1 through v3.5 VPred — each release pushing SDXL architecture closer to its theoretical ceiling while maintaining full backward compatibility.

Core Differentiators
  • Native 1536px Resolution: Zero upscaler dependency, edge-sharp output directly from the model
  • Hybrid Prompting: Both natural language and Danbooru tags, offering tag precision + NLP compositional understanding
  • Untuned Base Canvas: No baked-in aesthetic biases, the cleanest starting point for LoRA and fine-tuning
  • Iteration History: v0.1 → v1.0 → v1.1 → v2.0 STABLE → v3.0 → v3.5 VPred, representing the complete evolution of SDXL anime generation

Illustrious XL Core Features

1. Native 1536px High-Resolution Generation

Standard SDXL models are trained at 1024×1024, requiring Hires.fix or external upscalers — both prone to artifacts — to reach higher resolutions. Illustrious XL breaks this ceiling by training directly at 1536×1536, delivering native high-resolution output without post-processing amplification. The model also supports non-standard aspect ratios such as 1248×1824 without duplicating subjects or breaking composition, preserving complex layouts at their intended proportions. Edge regions remain sharp; fine linework and intricate costume details render without the blur or distortion typical of upscaled generations.

2. Hybrid Prompting System (NLP + Danbooru Tags)

The hybrid prompting system distinguishes Illustrious XL from every other SDXL derivative. The model simultaneously understands natural language semantics (descriptive English sentences) and structured Danbooru tag semantics, automatically reconciling both input types during inference. Trained on the Danbooru2023 dataset with knowledge cutoffs through June 2024, it recognizes thousands of specific anime characters and multiple art styles — cel shading, thick paint, 3D rendering.

Workflow example: Write "A girl sitting in a cafe, soft morning light, intricate lace dress" first, then append "1girl, red hair, sunlight, lace_trim, detailed_background". The model fuses the compositional understanding of the NLP description with the precision of Danbooru tags, achieving results neither input mode alone could produce.

  • Precision + Flexibility: Tags deliver exact character/attribute control; NLP handles scene composition, lighting, and atmosphere
  • Broad Knowledge Base: Trained on Danbooru2023, covering thousands of known characters and art styles up to June 2024
  • Reduced Prompt Engineering: Users can start with natural descriptions and refine with tags rather than memorizing tag syntax
  • Learning Curve for New Users: Understanding Danbooru tag conventions requires initial familiarization
  • Tag-NLP Conflicts: Occasionally, strong tag signals can override NLP-driven composition, requiring prompt tuning

3. Stable Text Encoder Fine-Tuning

Conventional SDXL training freezes the text encoder to maintain training stability, but this also limits concept understanding. OnomaAI Research developed a proprietary Stable Text Encoder Fine-Tuning technique that breaks this convention, enabling deeper comprehension of complex actions, multi-character compositions, and cross-subject prompts. The result is significantly reduced prompt bleed — in multi-character scenarios, each character's attributes remain cleanly separated without cross-contamination.

4. V-Prediction Architecture (v3.5 VPred)

The latest major iteration, v3.5 VPred, replaces the standard Epsilon prediction (noise prediction) with V-Prediction (velocity prediction). This architectural shift fundamentally changes how the diffusion model handles contrast, dynamic range, and convergence speed. In practice, v3.5 VPred delivers deeper blacks, brighter highlights, and richer mid-tones — particularly valuable for high-contrast scenes, dramatic lighting, and complex perspective compositions. It represents the technical ceiling of the SDXL architecture.

  • Superior Contrast: Deeper blacks and brighter highlights compared to Epsilon-based models
  • Faster Convergence: V-Prediction achieves stable results in fewer sampling steps
  • Better Composition Stability: Complex multi-element scenes maintain structural integrity
  • Harder to Fine-Tune: V-Prediction requires different training approaches; not all existing LoRAs are directly compatible
  • Newer, Less Tested: The community has more tooling and experience with Epsilon-based v2.0 STABLE

5. Untuned Base Canvas (v2.0 STABLE)

Most SDXL models bake in strong aesthetic preferences — specific lighting styles, color palettes, or rendering approaches. While this benefits out-of-the-box output, it contaminates LoRA training by introducing unavoidable stylistic interference. Illustrious XL's v2.0 STABLE deliberately avoids this: it is intentionally not tuned for any specific aesthetic, making it the "zero point" for customization. The model deeply understands concepts without forcing stylistic outputs, achieving high concept separation. This is why v2.0 STABLE has become the community's preferred parent model for training custom checkpoints and LoRAs on Civitai.

6. Seamless UI Ecosystem Integration

The model ships with full compatibility across the major AI generation toolchains: ComfyUI (with official custom nodes), Automatic1111, Forge, and Draw Things. OnomaAI provides an official VAE and dedicated ComfyUI_Illustrious custom nodes via GitHub, enabling predictable generation behavior across environments — critical for automated pipelines and workflow sharing.


Who Uses Illustrious XL

Freelance Illustrators Creating High-Resolution Anime Art

Professional illustrators working with AI tools face a persistent bottleneck: generating large-form artwork requires post-processing magnification, which degrades line precision and introduces softness. Illustrious XL's native 1536px output eliminates this dependency. Artists report eliminating Hires.fix from their workflow entirely, producing portfolio-ready images with sharp lines and rich detail in a single generation pass. User ratings on the official site average 5/5 stars, with reviewers specifically highlighting the resolution advantage as transformative for professional output.

LoRA Trainers and Model Fine-Tuners

The LoRA training community has a well-known problem: most parent models carry baked-in aesthetic biases that bleed into trained LoRAs, creating style interference when combining multiple LoRAs in a single generation. Illustrious XL's v1.0 and v2.0 STABLE checkpoints solve this by design — their untuned canvas architecture provides the cleanest training foundation available. Trained LoRAs exhibit higher concept separation, cleaner style independence, and better composability when used alongside other LoRAs.

💡 Version Selection for LoRA Training

For LoRA training, v2.0 STABLE offers the most predictable behavior thanks to its cosine annealing scheduler and extensive community validation. For one-shot generation quality, v3.5 VPred delivers superior contrast and dynamic range. Choose based on your primary workflow: training stability → v2.0 STABLE, output quality → v3.5 VPred.

Workflow Developers and Automation Pipeline Builders

Developers building automated generation workflows need predictable, reproducible behavior across tools. Illustrious XL's full compatibility with ComfyUI, A1111, Forge, and Draw Things — combined with official VAE and custom nodes — ensures consistent output regardless of the inference environment. The Hybrid Prompting system adds another layer of determinism: structured Danbooru tags provide precise control parameters, while NLP handles scene-level composition, enabling repeatable results at scale.

Researchers and Frontier Architecture Explorers

Research teams tracking the evolution of AI illustration architecture find Illustrious XL uniquely positioned. It represents both the culmination of SDXL architecture (v3.5 VPred pushes it to its technical limit) and a bridge to the next generation — the transition to Illustrious LU (Lumina Image architecture) with early v0.03/v1.0 weights already available. This dual role makes ILXL a living case study in architectural evolution, directly accessible to the research community.


Technical Architecture

Native Resolution Breakthrough

Standard SDXL models are constrained to 1024px native resolution — a direct limitation of the architecture's training regime. Illustrious XL breaks this by training at 1536×1536, operating at the SDXL framework's theoretical resolution ceiling. The model supports non-standard aspect ratios (e.g., 1248×1824) without subject duplication or composition breakage, a common failure mode in upscaled outputs. Edge regions maintain sharpness without the blur artifacts introduced by post-processing amplifiers. This is not an upscaling trick — it is native-resolution inference.

V-Prediction vs. Epsilon Prediction

The shift from Epsilon (v2.0 STABLE) to V-Prediction (v3.5 VPred) represents a fundamental change in the diffusion model's optimization target:

Parameter v2.0 STABLE (Epsilon) v3.5 VPred (Velocity)
Prediction Target Noise (epsilon) Velocity (v)
Contrast Handling Standard Superior — deeper blacks, brighter highlights
Convergence Speed Standard Faster — stable output in fewer steps
Best For LoRA training, predictable output One-shot quality, high-contrast scenes
Compatibility Broad LoRA ecosystem Newer, partial LoRA compatibility
Architecture Stage Stable, community-validated Technical ceiling of SDXL
  • v2.0 STABLE: Most predictable generation behavior, broadest LoRA compatibility, ideal training foundation
  • v3.5 VPred: Superior dynamic range, faster convergence, higher visual quality ceiling
  • v2.0 STABLE: Lower contrast ceiling, requires more sampling steps for equivalent detail
  • v3.5 VPred: Smaller compatible LoRA ecosystem, different training requirements for fine-tuning

Stable Text Encoder Fine-Tuning

OnomaAI's proprietary fine-tuning technique departs from the industry convention of freezing the text encoder during training. By enabling stable, controlled updates to the text encoder weights, the model achieves deeper understanding of complex action instructions, multi-character interactions, and cross-subject attribute assignments. The primary measurable benefit is reduced prompt bleed — in generations with multiple characters, each subject maintains its defined attributes without attribute leakage between characters.

Cosine Annealing Learning Rate Scheduling

Introduced in v2.0 STABLE, this scheduler replaces aggressive learning rate decay with a smooth cosine curve. The practical impact is significantly reduced unpredictable generation behavior — output variance across identical seeds and prompts drops measurably compared to earlier versions. This is the technical foundation for v2.0 STABLE's reputation as the most reliable training base.

Danbooru2023 Training Dataset

The training corpus draws from the Danbooru2023 dataset, providing comprehensive coverage of anime characters, art styles, and technical rendering approaches — cel shading, thick paint, 3D rendering, and hybrid techniques. Knowledge cutoff is June 2024, giving the model contemporary character and style recognition. This dataset breadth is what enables the Hybrid Prompting system's recognition of thousands of specific named characters and stylistic references.


Ecosystem and Integration

Supported Platforms and Frameworks

Illustrious XL is designed for drop-in integration across the dominant AI generation ecosystems:

Platform Compatibility Notes
ComfyUI Full Official custom nodes via GitHub (ComfyUI_Illustrious)
Automatic1111 Full Standard checkpoint loading, no additional setup
Forge Full Fully compatible with Forge's optimizations
Draw Things Full iOS inference support

The official VAE ensures color reproduction consistency across all platforms. Generation behavior remains highly predictable across environments — a critical requirement for developers building multi-platform pipelines.

Model Distribution and Community Ecosystem

Primary model weights are distributed via Hugging Face (OnomaAIResearch/Illustrious-XL-v2.0), with additional checkpoints and variant weights available through the official release pipeline. The community ecosystem on Civitai is extensive: thousands of derivative models and compatible LoRAs list ILXL checkpoints as their parent model. This creates a self-reinforcing ecosystem where new LoRA releases expand the model's capabilities, which in turn attracts more users to the platform.

Online Experience

Users can test Illustrious XL immediately without local setup through:

  • Official Playground at illustriousxl.org/playground
  • WAI Illustrious service at illustriousxl.org/wai-illustrious-sdxl
  • Pollo AI (associated platform)

These provide immediate access to the v3.5 VPred model for evaluation and one-off generation without GPU requirements.

Next-Generation Architecture: Illustrious LU

The SDXL architecture, pushed to its technical limit by v3.5 VPred, has a successor in development. Illustrious LU is built on the Lumina Image architecture, designed to overcome the scaling limitations inherent in SDXL. Early weights — v0.03 and v1.0 — are already available, enabling researchers and early adopters to begin experimentation.

💡 Architecture Transition Timeline

The SDXL architecture has reached its practical ceiling with v3.5 VPred. Illustrious LU (Lumina Image) represents the architectural evolution path. Early weights (v0.03/v1.0) are available for preview. Users running production workflows should continue with v2.0 STABLE or v3.5 VPred for stability while monitoring LU releases for architectural maturity.

Open-Source Licensing

Illustrious XL is distributed under the CreativeML OpenRAIL-M license combined with the Fair AI Public License. Key terms:

  • Free for personal and research use — no licensing cost
  • Derivative models and LoRAs permitted — community fine-tuning explicitly encouraged
  • Proprietary closed-source commercialization prohibited — derivative works must remain open

This licensing structure balances open community development with protection against proprietary enclosure, consistent with the project's mission as community infrastructure.


Frequently Asked Questions

What is the essential difference between Illustrious XL and standard SDXL models?

Three fundamental differences: (1) Native 1536px resolution — trained at 1536×1536, not upscaled, so output is sharp without Hires.fix; (2) Hybrid Prompting — accepts both natural language and Danbooru tags simultaneously; (3) Untuned base canvas — deliberately free of aesthetic bias, making it the cleanest foundation for LoRA and custom checkpoint training. Standard SDXL models typically cap at 1024px, accept only natural language, and carry baked-in aesthetic preferences.

Which version should I use: v2.0 STABLE or v3.5 VPred?

It depends on your primary workflow. Choose v2.0 STABLE if you are training LoRAs or custom checkpoints — its cosine annealing scheduler provides the most predictable generation behavior, and it has the broadest ecosystem of compatible community LoRAs. Choose v3.5 VPred if your priority is one-shot generation quality — its V-Prediction architecture delivers superior contrast, deeper blacks, brighter highlights, and faster convergence. v3.5 VPred represents the SDXL architecture's technical ceiling, but its LoRA ecosystem is smaller and its training characteristics differ from v2.0 STABLE.

Which tools and platforms does Illustrious XL support?

The model is fully compatible with ComfyUI (official custom nodes available at github.com/onomaai/ComfyUI_Illustrious), Automatic1111, Forge, and Draw Things. OnomaAI provides an official VAE for consistent color reproduction and the ComfyUI_Illustrious custom nodes for advanced workflow integration. Model weights are distributed via Hugging Face at OnomaAIResearch/Illustrious-XL-v2.0.

How do I write effective hybrid prompts?

Start with a natural language scene description — for example, "A girl in a white dress standing in a field of sunflowers, golden hour lighting, soft bokeh". Then append Danbooru-style tags separated by commas: "1girl, white_dress, sunflowers, golden_hour, bokeh, intricate_details, masterpiece". The model automatically fuses both inputs — the NLP provides compositional and atmospheric understanding, while the tags enforce precise attribute control. Experienced users commonly iterate: write the scene in NLP, then refine with tags for specific character attributes, clothing details, and style references.

What are the licensing restrictions for Illustrious XL?

Illustrious XL is distributed under the CreativeML OpenRAIL-M license and the Fair AI Public License. It is free for personal and research use. Derivative models, LoRAs, and fine-tuned checkpoints are permitted and encouraged. However, proprietary closed-source commercialization is prohibited — any commercial use must comply with the open licensing terms. Full license details are available at the official website. This licensing is designed to protect community access while preventing proprietary enclosure.

What is Illustrious LU and how does it relate to Illustrious XL?

Illustrious LU is OnomaAI Research's next-generation model built on the Lumina Image architecture. It represents the successor to the SDXL-based Illustrious XL line. Because SDXL architecture has been pushed to its technical limit with v3.5 VPred, the team has transitioned development to the more scalable Lumina architecture. Early weights — v0.03 and v1.0 — are already available for preview. Users should view Illustrious LU as the architectural evolution path while continuing production workflows on the stable SDXL-based releases.

Where can I download Illustrious XL models?

Primary model weights are available on Hugging Face at huggingface.co/OnomaAIResearch/Illustrious-XL-v2.0. For online testing without local setup, use the official Playground at illustriousxl.org/playground or the WAI Illustrious service at illustriousxl.org/wai-illustrious-sdxl. The associated Pollo AI platform also provides access. Custom nodes for ComfyUI are available via GitHub at github.com/onomaai/ComfyUI_Illustrious.

Comments

Comments

Please sign in to leave a comment.
No comments yet. Be the first to share your thoughts!