🧂 Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

Preprint
1The Hong Kong University of Science and Technology   2Vivix Group Limited   3The University of Sydney   4Westlake University
† Project Lead  ·  ‡ Corresponding Author  ·  xingtong.ge@gmail.com
LTX-2 · bidirectional · 40 steps
OmniForcing · causal · 4 steps
Salt++ (TF-dCM route) · 4 steps
Salt++ (AR–AR DMD route) · 4 steps
LTX-2 · bidirectional · 40 steps
OmniForcing · causal · 4 steps
Salt++ (TF-dCM route) · 4 steps
Salt++ (AR–AR DMD route) · 4 steps
01 / 02

Fast motion at the same four-step budget

Same prompt and seed. LTX-2 is bidirectional at 40 steps; OmniForcing and both Salt++ routes are causal at four steps. Tap a panel's ♫ chip to switch which soundtrack you hear.

01 / Overview

Abstract

Few-step streaming audio–video generation requires both causal modeling and step distillation. Teacher forcing pairs clean history with a noisy target, but shapes predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student predicts intermediate representations from a clean-history exponential-moving-average teacher. Together with flow matching, this self-supervised signal encourages the student to extract semantic information useful for next-block prediction and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A scale-wise post-training stage further extends Salt++ to 4-step 1664×960 generation, outperforming bidirectional LTX-2 on six of seven reported metrics.

Joint audio and video

A single causal generator produces temporally aligned video and audio blocks, so every clip below carries its own generated soundtrack.

Causal Self-Flow

Varying only the history while holding the noisy target fixed turns information asymmetry into a representation-learning signal for the autoregressive teacher.

Context-aligned AR DMD

Generator sampling and both score models share one causal mask and prefix, so distribution matching alone suffices — no separate consistency-distillation stage.

02 / Method

Method

Salt++ keeps one autoregressive DMD objective throughout post-training and only shifts its conditioning from clean context to generated rollout.

Comparison with recent causal post-training recipes

Prior recipes switch objectives between stages and score the causal generator with bidirectional models. Salt++ keeps a single AR DMD objective; a further scale-wise stage reaches 1664×960 with four generator calls.

Two roles of causal context in post-training

Two roles of causal context. Left: under teacher forcing the clean history carries information the flow objective never supervises directly; Causal Self-Flow turns that asymmetry into representation supervision. Right: each model's visibility over block positions — context alignment makes the three masks identical.

Score-context configurations in causal DMD

Score–context configurations. A bidirectional score reads the whole clip at one noise level; an autoregressive score reads the clean prefix without future access. Only AR–AR scores each block under the context that produced it.

Cross-modal alignment during AR teacher training

Causal Self-Flow develops stronger audio–visual and audio–text alignment than clean teacher forcing throughout AR teacher training.

03 / Videos

Generated results

Audio and video are generated jointly. Every clip carries its own generated soundtrack, so audio–visual synchrony can be judged directly. Press Unmute, or Play together to start a whole row in sync.

480p qualitative comparison between LTX-2 Base, OmniForcing and both Salt++ routes

Four-step streaming audio–video generation at 480p. Salt++ retains scene and facial detail that the causal baseline loses at the same step budget.

480p streaming generation

Four-step causal generation against the bidirectional foundation model and the causal baseline, under matched prompts and seeds.

Wellness host · paper Fig. 8
[Shot 1][0.0s-5.0s]: A calm wellness-channel video in a sunlit living room. A Caucasian man in his late 60s with short white hair and a neatly trimmed white beard sits upright in a simple armchair. A static chest-up composition and soft frontal daylight keep his face natural, unobstructed, and sharply focused. He wears a pale blue shirt. Both hands rest separately on the armrests and remain fully visible. He looks directly into the lens, takes a small breath, and says in a warm, steady voice, "Rest is not a reward for finishing everything. Nothing is ever finished." He relaxes his shoulders slightly and gives one gentle nod. The audio is his voice, a distant bird outside the window, and quiet room ambience; the camera never moves.
LTX-2 Base
40 steps
OmniForcing
4 steps
Salt++ (TF-dCM route)
4 steps
Salt++ (AR–AR DMD route)
4 steps

1664×960 scale-wise generation

Two low-resolution and two high-resolution generator calls per block, still four calls in total.

Fitness coach · paper Fig. 6
[Shot 1][0.0s-5.0s]: In a minimalist fitness studio with a neutral grey wall, a mature Black man with a shaved head and a short grey beard stands on a black mat. The camera is locked in an eye-level medium shot, and broad soft lighting keeps his face, shoulders, and hands crisp without motion blur. He wears a maroon athletic shirt. Both hands begin open at waist height, palms angled inward and clearly separated. He brings them slowly upward by a few inches while maintaining direct eye contact and says in a calm, supportive voice, "You don't need more motivation; you need one smaller promise." His hands stop and remain still as he gives a gentle nod. The studio is quiet except for his resonant voice and faint ventilation hum.
LTX-2 Base
40 + 3 steps
OmniForcing
4 steps, direct 960p
Salt++ 480p
4 steps, bilinear 2×
Salt++ 960p
2 LR + 2 HR

Ablations

Score–context configuration for few-step initialization, and teacher-guidance calibration during distillation.

Cherry-blossom garden · paper Fig. 7
A serene garden with cherry blossom trees featuring pink and white flowers is shown. In the background, a traditional Japanese-style building with a sloped roof stands. The garden includes a pond with clear water surrounded by rocks and greenery. The sky is clear and bright, indicating a sunny day. Soft, gentle piano music plays in the background throughout.
TF-dCM
consistency distillation
BI–BI DMD
both scores bidirectional
BI–AR DMD
fake score causal
AR–AR DMD
context aligned
04 / Results

Quantitative results

JavisBench-mini and VBench, all causal models at four steps with CFG = 1.

+57%
Visual quality over OmniForcing, 480p, 4 steps
+45%
Motion quality over OmniForcing, 480p, 4 steps
1664×960
Four generator calls per block, scale-wise
Table 1(a). JavisBench-mini, official prompts, 480p
ModelCausalStepsVQ↑MQ↑AQ↑CLIP↑IB-AV↑Javis↑DeSync↓
LTX-2 BaseNo401.8840.5664.9860.3110.2390.2000.608
AR teacher (CSF)Yes402.3040.8574.5650.3160.2020.1620.746
OmniForcingYes41.8070.6994.7180.3030.1630.1240.745
Salt++ (TF-dCM route)Yes42.0130.8264.9760.3130.2290.1850.710
Salt++ (AR–AR DMD route)Yes42.8381.0104.9910.3160.1840.1460.759

Bold is best and underline is second best among the 4-step rows.

Table 1(b). JavisBench-mini, official prompts, 1664×960
ModelCausalStepsVQ↑MQ↑AQ↑CLIP↑IB-AV↑Javis↑DeSync↓
LTX-2 BaseNo40+32.2310.6074.8670.3110.1690.1450.658
OmniForcingYes42.2980.8324.7210.2930.1650.1280.755
Salt++Yes42.7300.9575.1130.3180.1930.1570.768

Bold marks the best result across all methods.

Table 1(c). VBench video quality and within-clip consistency, 480p
ModelCausalStepsAesthetic↑Imaging↑Subject↑Background↑
LTX-2 BaseNo4053.8967.0695.9495.63
OmniForcingYes456.7468.4196.6295.32
Salt++ (TF-dCM route)Yes453.6562.8796.5896.05
Salt++ (AR–AR DMD route)Yes455.4369.9897.0196.24
Table 2. 1,000 LTX-2-enhanced JavisBench prompts, 480p
ModelCausalStepsVQ↑MQ↑AQ↑IB-AV↑AVH↑Javis↑DeSync↓
LTX-2 BaseNo401.9810.6404.9590.2560.2450.2110.573
OmniForcingYes41.7740.6074.6130.1580.1520.1210.734
Salt++ (TF-dCM route)Yes42.0740.7604.8380.2450.2330.1930.704
Salt++ (AR–AR DMD route)Yes42.8200.9994.9970.1920.1860.1510.696
05 / Ablations

Ablations

Which conditional distributions are compared, and which reference is distilled.

Table 3(a). Few-step initialization: score contexts and objectives
ObjectiveScoresVQ↑MQ↑AQ↑CLIP↑IB-TV↑IB-TA↑DeSync↓
TF-dCM–1.9160.6114.6860.3070.2640.1450.766
TF-DMDBI/BI0.8370.1404.7960.3010.2660.1440.811
TF-DMDBI/AR1.1310.1574.5710.2870.2550.1490.531
TF-DMDAR/AR3.1471.3514.7500.3170.2710.1530.726

BI/AR denotes the visibility of the real and fake score models respectively.

Table 3(b). Teacher guidance under aligned AR–AR scores
Teacher guidanceVQ↑ (off.)MQ↑ (off.)VQ↑ (rew.)MQ↑ (rew.)AQ↑CLIP↑DeSync↓
Fixed 4.01.6430.5721.7120.5984.9340.3150.843
Randomized U(1.0, 3.5)3.1471.3513.1321.2614.7500.3170.726

Guidance inherited from consistency distillation versus guidance calibrated for DMD.

06 / BibTeX

BibTeX

@misc{ge2026saltcontextalignedposttrainingfewstep,
  title     = {Salt++: Context-Aligned Post-Training for Few-Step Streaming
               Multimodal Generation},
  author    = {Xingtong Ge and Yutong Wang and Lunjie Zhu and Haitao Lin and
               Fangyu Lin and Yushi Huang and Xin Zhang and Yi Zhang and
               Yu Liu and Jun Zhang},
  year      = {2026},
  eprint    = {2609.36995},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url       = {https://arxiv.org/abs/2609.36995}
}