AKUSPACE: Acoustic Space LoRA for LTX-2.5 Audio
KoshiMazaki releases AKUSPACE v0.5, a spatial audio LoRA for LTX-2.5 that places generated voices and music in rooms, clubs, cathedrals or outdoors, with ComfyUI nodes included.
KoshiMazaki released AKUSPACE v0.5, a spatially aware audio LoRA for LTX-2.5 that gives generated or reference audio the acoustic character of a physical space: a small room, an empty club, a cathedral, or an outdoor environment. The LoRA, demo site, and a companion ComfyUI node pack are all open: Hugging Face | interactive demo | ComfyUI-Koshi-Nodes.
The AKUSPACE control surface: source, listener and the selected acoustic space, with the trained decay time read out beneath.
What it does
AKUSPACE is trained as an audio-to-audio adapter: it transforms audio it is given, shaping voices, beats and instruments with room reverb, outdoor ambience and spatial sound effects. The control surface is prompt-driven, with the trigger word AKUSPACE and a level word between the space and its character:
AKUSPACE female spoken voice through synthetic cathedral reverb, moderate wide diffuse reflections and a long decaying tail, no background ambience| Mode | Options | Levels |
|---|---|---|
| Space: rooms a sound sits in | small room, medium room, empty club, cathedral | gentle / moderate / heavy |
| Place: environments a sound sits among | outdoor day, outdoor night | gentle / heavy |
| Sound effects: processing a sound goes through | dual delay | gentle / heavy |
Outdoor "places" only scale down, not up: an ambience bed is a separate recording rather than a reverb tail. Room captions carry a decay time; cathedral, the granular effect and both outdoor places have no numeric decay.
Supported workflows
| Input | Output |
|---|---|
| Audio | Audio-to-audio treatment for an existing recording |
| Text + audio | Text-to-video with AKUSPACE-treated synchronized audio |
| Image + audio | Image-to-video with AKUSPACE-treated synchronized audio |
The proven route for video is dry voice, then an audio-to-audio space pass, then image+audio-to-video with the treated audio held fixed while the image conditions the first frame. The adapter targets the audio branches; video generation is handled by the base LTX-2.5 model. It can also run as a dubbing pass, where its time-aligned reference keeps the treatment anchored to the source performance.
Settings and usage
Suggested settings are 24 steps, CFG 1-2, adjusting to CFG 4 for higher volume and detail. Partial captions work less well than complete ones, since the trailing clauses were present in every training caption. The ComfyUI node pack ships a three.js viewer for manipulating the prompt from a scene view: source, listener and space.
Availability
- LoRA weights: KoshiMazaki/akuspace-ltx25 under the LTX-2 community license
- ComfyUI nodes: koshimazaki/ComfyUI-Koshi-Nodes
- Interactive audiovisual study: akuspace.pages.dev
Comments
Sign in with GitHub to join the discussion.