AKUSPACE: Acoustic Space LoRA for LTX-2.5 Audio

ComfyUI Wikinews

KoshiMazaki releases AKUSPACE v0.5, a spatial audio LoRA for LTX-2.5 that places generated voices and music in rooms, clubs, cathedrals or outdoors, with ComfyUI nodes included.

KoshiMazaki released AKUSPACE v0.5, a spatially aware audio LoRA for LTX-2.5 that gives generated or reference audio the acoustic character of a physical space: a small room, an empty club, a cathedral, or an outdoor environment. The LoRA, demo site, and a companion ComfyUI node pack are all open: Hugging Face | interactive demo | ComfyUI-Koshi-Nodes.

AKUSPACE control surface

The AKUSPACE control surface: source, listener and the selected acoustic space, with the trained decay time read out beneath.

What it does

AKUSPACE is trained as an audio-to-audio adapter: it transforms audio it is given, shaping voices, beats and instruments with room reverb, outdoor ambience and spatial sound effects. The control surface is prompt-driven, with the trigger word AKUSPACE and a level word between the space and its character:

AKUSPACE female spoken voice through synthetic cathedral reverb, moderate wide diffuse reflections and a long decaying tail, no background ambience
ModeOptionsLevels
Space: rooms a sound sits insmall room, medium room, empty club, cathedralgentle / moderate / heavy
Place: environments a sound sits amongoutdoor day, outdoor nightgentle / heavy
Sound effects: processing a sound goes throughdual delaygentle / heavy

Outdoor "places" only scale down, not up: an ambience bed is a separate recording rather than a reverb tail. Room captions carry a decay time; cathedral, the granular effect and both outdoor places have no numeric decay.

Supported workflows

InputOutput
AudioAudio-to-audio treatment for an existing recording
Text + audioText-to-video with AKUSPACE-treated synchronized audio
Image + audioImage-to-video with AKUSPACE-treated synchronized audio

The proven route for video is dry voice, then an audio-to-audio space pass, then image+audio-to-video with the treated audio held fixed while the image conditions the first frame. The adapter targets the audio branches; video generation is handled by the base LTX-2.5 model. It can also run as a dubbing pass, where its time-aligned reference keeps the treatment anchored to the source performance.

Turn prompt enhancement off. The rewriter paraphrases away the trigger word and the level word, which are exactly the tokens the control surface depends on. Also note that with the LoRA loaded and no reference audio present, text-to-audio produces near-silence: there is nothing to transform.

Settings and usage

Suggested settings are 24 steps, CFG 1-2, adjusting to CFG 4 for higher volume and detail. Partial captions work less well than complete ones, since the trailing clauses were present in every training caption. The ComfyUI node pack ships a three.js viewer for manipulating the prompt from a scene view: source, listener and space.

Availability

Comments

Sign in with GitHub to join the discussion.

Loading comments…
AKUSPACE: Acoustic Space LoRA for LTX-2.5 Audio | ComfyUI Wiki