MambaFlow-TTS

Eliminating Multiscale Aliasing in Flow-Matching Speech Synthesis via Full-Resolution State-Space Models

Developed by Monesh

Welcome to the Demonstration

Hello and welcome to the MambaFlow Text-to-Speech audio demonstration site. This research focuses on eliminating severe temporal aliasing—commonly known as comb-filtering—which plagues modern diffusion and flow-matching speech systems. Through rigorous ablation studies, we prove that multiscale downsampling and transposed convolutions are the root cause of these robotic artifacts. Our proposed champion architecture, the Tetra Sequential Decoder, operates entirely at a 1x native resolution, achieving high spectral convergence. Please enjoy the comparative audio samples below, and pay special attention to the pristine high-frequency clarity of the Tetra model.

Champion Model: Tetra Sequential Decoder (1x Native Res) Zero Aliasing

Architectures Evaluated

Our benchmark rigorously tests 5 unique architectural topologies to observe the effects of multiscale processing on acoustic clarity.

1. Tetra Sequential (Ours) 🏆

Operates entirely at 1x native resolution using Bidirectional Mamba-2 blocks. Eliminates multiscale aliasing completely and avoids transposed convolutions.

2. Staircase-XTEncoder

A multi-scale feature pyramid (1x → 1/2x → 1/4x) using cross-attention local smoothing. Achieves very low L1 errors but introduces minor phase artifacts.

3. Staircase-Mamba2

A ResNet-style feature pyramid using purely Bidirectional Mamba-2 and ConvNeXt blocks with lateral skip connections.

4. TwoStage-Mamba2

Cascaded dual-stage downsampling (1/2x and 1/4x) which relies heavily on multiscale transposed convolutions, producing audible "dual voice" flanging.

5. UNet-OneBottleneck

A standard UNet topology focused on high throughput and memory efficiency, serving as the fast but artifact-prone baseline.

Ablation Study: Comparative Audio Samples

Listen to the subtle differences in phase and temporal aliasing across the 5 architectural topologies. The robotic "dual voice" effect is most prominent in the TwoStage and UNet baselines due to multiscale upsampling.