Anima-Light-Lavender: Anima Post-Train for Long Natural Prompts
Anima-Light-Lavender is a post-train of Anima-Base v1.0 that reads 512-token natural-language captions, ships BF16 and MXFP8 weights and drops into ComfyUI unchanged.
Anima-Base v1.0 takes Danbooru-style tags as its native input. Anima-Light-Lavender moves in the other direction: it is trained on structured captions with a long natural-language image_description field, and the release focuses on understanding descriptions at the 512-token scale. The training set is roughly 1.7M Danbooru images, annotated and filtered by a data pipeline whose annotation, filtering and preference-judgment models were all trained by the author rather than borrowed off the shelf.
The model targets anime illustrations, characters and stylized art, and is driven by descriptive passages rather than tag stacks.
What changed, and what did not
| Anima-Light-Lavender | |
|---|---|
| Base model | Anima-Base v1.0 (anima-base-v1.0.safetensors) |
| Architecture | Unchanged: ~2B parameters, 28-layer DiT, no added layers and no distillation |
| Text encoder | Qwen3-0.6B (qwen_3_06b_base.safetensors) |
| VAE | Qwen-Image VAE (qwen_image_vae.safetensors) |
| Weights | anima-light-lavender.safetensors (BF16) and anima-light-lavender_mxfp8.safetensors (MXFP8, faster inference) |
| Training focus | Natural-language descriptions at the 512-token scale, plus higher-quality data |
Because the layer count, parameter count and file layout match the base model, an existing Anima workflow only needs its checkpoint swapped.
Installing it in ComfyUI
- Put one weight file in
ComfyUI/models/diffusion_models/. Pick either the BF16 or the MXFP8 version, not both. - Reuse the text encoder and VAE from your current Anima setup:
qwen_3_06b_base.safetensorsinComfyUI/models/text_encoders/,qwen_image_vae.safetensorsinComfyUI/models/vae/. - Install the companion nodes with
git clone https://github.com/aa0525/comfyui-zako-pe.gitinsideComfyUI/custom_nodes/, then restart ComfyUI.
The comfyui-zako-pe pack adds two nodes: Danbooru Caption JSON, which assembles the structured caption, and Danbooru Prompt Extend (OpenAI), which calls an OpenAI-compatible server to expand a tag-style image_description into natural language. The companion model zako-pe (ZAKO-V0.1) is designed to serve that expansion locally.
The release ships two workflows, both loadable by dragging the JSON onto the ComfyUI canvas:
Both default to the MXFP8 weights. If you installed the BF16 file instead, switch the checkpoint in the Load Diffusion Model node.
Prompting with a structured caption
The model is trained on a fixed-field caption, with the description written out as prose:
{
"year": 2025,
"preference_level": "best",
"artist": [...],
"copyright": [...],
"character": [...],
"image_description": "...",
"extra_tags": [...]
}| Field | Purpose |
|---|---|
year | Year anchor, default 2025 |
preference_level | normal, high, very_high or best (default best) |
artist | Artist and style tags |
copyright | Series or franchise tags |
character | Character tags |
image_description | The image content, written as a detailed natural-language passage |
extra_tags | Supplementary content tags |
If writing a long passage is not what you want, the basic workflow accepts tags and lets the prompt-extend node rewrite them. The author's own alternative is to describe the image and let the model carry the rest: the release notes that results stay clean even without quality words or a negative prompt.
Recommended settings
| Parameter | Value |
|---|---|
| Sampler | euler |
| Scheduler | simple |
| Steps | 25 |
| CFG | 4.0 |
| Resolution | Around 1280×1280 |
| Negative prompt | Leave empty |
Training details
Training ran in BF16 with MXFP8 GEMM in the MLP and attention matrix multiplications, a batch size of 1024 and a peak learning rate of 4e-5. MLP and attention 2D parameters were updated by Muon (momentum 0.95, match_rms_adamw) while the remaining parameters stayed on AdamW (β 0.9 / 0.95, ε 1e-8). Caption preference_level tiers come from a preference model trained on about 3M samples, and the training data runs to the end of November 2025.
Comments
Sign in with GitHub to join the discussion.