Anima-Light-Lavender: Anima Post-Train for Long Natural Prompts

ComfyUI Wikinews

Anima-Light-Lavender is a post-train of Anima-Base v1.0 that reads 512-token natural-language captions, ships BF16 and MXFP8 weights and drops into ComfyUI unchanged.

Anima-Light-Lavender (Hugging Face) is a post-train of Anima-Base v1.0, the 2B anime model from Circlestone Labs. Johnny-Z published it on September 18, and it keeps the base model's architecture intact, so it replaces the upstream weights file for file.

Anima-Base v1.0 takes Danbooru-style tags as its native input. Anima-Light-Lavender moves in the other direction: it is trained on structured captions with a long natural-language image_description field, and the release focuses on understanding descriptions at the 512-token scale. The training set is roughly 1.7M Danbooru images, annotated and filtered by a data pipeline whose annotation, filtering and preference-judgment models were all trained by the author rather than borrowed off the shelf.

Anima-Light-Lavender sample banner

The model targets anime illustrations, characters and stylized art, and is driven by descriptive passages rather than tag stacks.

What changed, and what did not

Anima-Light-Lavender
Base modelAnima-Base v1.0 (anima-base-v1.0.safetensors)
ArchitectureUnchanged: ~2B parameters, 28-layer DiT, no added layers and no distillation
Text encoderQwen3-0.6B (qwen_3_06b_base.safetensors)
VAEQwen-Image VAE (qwen_image_vae.safetensors)
Weightsanima-light-lavender.safetensors (BF16) and anima-light-lavender_mxfp8.safetensors (MXFP8, faster inference)
Training focusNatural-language descriptions at the 512-token scale, plus higher-quality data

Because the layer count, parameter count and file layout match the base model, an existing Anima workflow only needs its checkpoint swapped.

Installing it in ComfyUI

  1. Put one weight file in ComfyUI/models/diffusion_models/. Pick either the BF16 or the MXFP8 version, not both.
  2. Reuse the text encoder and VAE from your current Anima setup: qwen_3_06b_base.safetensors in ComfyUI/models/text_encoders/, qwen_image_vae.safetensors in ComfyUI/models/vae/.
  3. Install the companion nodes with git clone https://github.com/aa0525/comfyui-zako-pe.git inside ComfyUI/custom_nodes/, then restart ComfyUI.

The comfyui-zako-pe pack adds two nodes: Danbooru Caption JSON, which assembles the structured caption, and Danbooru Prompt Extend (OpenAI), which calls an OpenAI-compatible server to expand a tag-style image_description into natural language. The companion model zako-pe (ZAKO-V0.1) is designed to serve that expansion locally.

The release ships two workflows, both loadable by dragging the JSON onto the ComfyUI canvas:

Both default to the MXFP8 weights. If you installed the BF16 file instead, switch the checkpoint in the Load Diffusion Model node.

Prompting with a structured caption

The model is trained on a fixed-field caption, with the description written out as prose:

{
  "year": 2025,
  "preference_level": "best",
  "artist": [...],
  "copyright": [...],
  "character": [...],
  "image_description": "...",
  "extra_tags": [...]
}
FieldPurpose
yearYear anchor, default 2025
preference_levelnormal, high, very_high or best (default best)
artistArtist and style tags
copyrightSeries or franchise tags
characterCharacter tags
image_descriptionThe image content, written as a detailed natural-language passage
extra_tagsSupplementary content tags

If writing a long passage is not what you want, the basic workflow accepts tags and lets the prompt-extend node rewrite them. The author's own alternative is to describe the image and let the model carry the rest: the release notes that results stay clean even without quality words or a negative prompt.

ParameterValue
Samplereuler
Schedulersimple
Steps25
CFG4.0
ResolutionAround 1280×1280
Negative promptLeave empty

Training details

Training ran in BF16 with MXFP8 GEMM in the MLP and attention matrix multiplications, a batch size of 1024 and a peak learning rate of 4e-5. MLP and attention 2D parameters were updated by Muon (momentum 0.95, match_rms_adamw) while the remaining parameters stayed on AdamW (β 0.9 / 0.95, ε 1e-8). Caption preference_level tiers come from a preference model trained on about 3M samples, and the training data runs to the end of November 2025.

Where it does not fit: photorealism is out of scope for the Anima series, and long text rendering (signs, subtitles, multi-line sentences) is unreliable. Single words and short phrases usually come out correctly.

Comments

Sign in with GitHub to join the discussion.

Loading comments…
Anima-Light-Lavender: Anima Post-Train for Long Natural Prompts | ComfyUI Wiki