Research Paper Author

The Sequential Learning Group

NeurIPS 2026

NeurIPS 2026

Closing the Loop: Co-Evolving EM for Irregular Time Series Generation in Lifted Representations

Gal Fadlon*, Idan Arbiv*, Omri Azencot

Irregular time-series generation is usually solved open-loop: impute once, freeze, train. We show the binding constraint is iteration, not imputer design, and close the imputer–generator loop with a co-evolving Monte Carlo EM in lifted image space.

PapersoonarXivsoonCodesoon

Abstract

Authors

Gal Fadlon*🐪 Ben Gurion University of the Negev
Idan Arbiv*🐪 Ben Gurion University of the Negev
Omri Azencot🐪 Ben Gurion University of the Negev

*equal contribution

At a Glance

59%
Discriminative
15%
Predictive
73%
Context-FID
71%
Correlation

Average relative improvement over the strongest baseline, across ten datasets and four corruption regimes.

Abstract

Regular time-series generation from irregular time series is typically reduced to one-shot imputation: complete each partial sequence once, freeze the completions, and train a generator on the resulting surrogate dataset. We show that this open-loop strategy is structurally flawed. Point imputers collapse ambiguous missing regions to conditional averages, while stochastic imputers become stale as the generator evolves; the binding constraint is therefore iteration, not imputer design. We introduce a co-evolving Monte Carlo EM framework that closes the imputer–generator loop. Each E-step samples missing values from the posterior of the current generator, and each M-step retrains the generator on these completions, preserving unconditional generator training. Since our diffusion prior operates in a lifted 2D representation while observations live in time-series space, we further introduce Posterior Sampling for Lifted Representations (PSLR), which enforces operator-consistent conditioning, uncertainty preservation, and manifold consistency during the E-step. Across ten datasets and four corruption regimes, our method achieves state-of-the-art performance, improving over the strongest baselines by 59% in discriminative score, 15% in predictive score, 73% in Context-FID, and 71% in feature-correlation error. Ablations show that our approach reduces average discriminative score by 73% over single-step baselines, closing most of the remaining gap to the clean-data oracle.

Key Contributions

  1. 1

    Iteration is the binding constraint

    We show, theoretically and across a 10-dataset benchmark, that closed-loop training via co-evolving EM outperforms all published open-loop baselines. Iteration, not imputer design or architectural sophistication, is the primary driver of generative fidelity under missingness.

  2. 2

    A closed-loop framework for irregular generation

    A co-evolving Monte Carlo EM that alternates between sampling missing values and retraining the generator. To make it work in lifted space we develop PSLR, a plug-and-play posterior sampler for cross-space constraints, with three structural conditions (C1–C3) any valid dual-space E-step must satisfy.

  3. 3

    State-of-the-art on four corruption regimes

    New SOTA on the standard ten-dataset benchmark, recovering near-oracle performance with rapid convergence, plus three new, harder benchmarks: per-cell asynchronous missingness, contiguous block gaps, and continuous-time irregular sampling.

Why Close the Loop?

Open-Loop Imputation Is Structurally Flawed

Prior methods (GT-GAN, KoVAE, ImagenI2R) share a commit-and-forget template: a pretrained imputer completes each irregular sequence once, the completions are frozen, and a generator is trained on them. Missing values are ambiguous, the same partial sequence admits many valid completions, but an MSE-trained imputer resolves the ambiguity by predicting the conditional mean. Across a dataset this systematically shrinks variability, so the generator learns a narrower distribution than the true one.

A better imputer does not fix this. A stochastic imputer restores diversity, but its conditional distribution is fixed before training; as the generator improves, the imputer falls out of sync. Training is unbiased only when missing values are sampled from the distribution the trained generator itself would produce, a self-referential requirement no pretrained imputer can satisfy. Closing the loop is what achieves it at every iteration.

Open loop vs. closed loop

Irregular data
partial observations y
repeat ~5 iterations
E-step (PSLR)
sample x ~ pθ(x | y) from the current generator
M-step
retrain the generator on fresh completions
Generator
unconditional, near-oracle

Each E-step samples completions from the current generator's posterior; each M-step continues training from the previous weights on them.

The Conditional-Mean Trap

Histograms comparing a bimodal true conditional with TST-like, NCDE-like and co-evolving EM imputations

Figure 2: With a bimodal conditional p(x₂ | x₁ ≈ 0), deterministic TST-like and NCDE-like imputers collapse missing values to the conditional mean, placing mass between the true modes. Co-evolving EM samples from the current generator's posterior and recovers both modes.

The binding constraint is iteration: imputations must be refined repeatedly using feedback from the generator they are used to train.

Lower Imputation Error ≠ Better Generation

A controlled 2D experiment where the true distribution is known: two interleaved spirals, with each coordinate removed with probability 0.6. The method with the lowest imputation MSE produces the worst samples, collapsing the spirals into a regression-line cross. Co-evolving EM has almost twice the MSE yet recovers both arms.

Generated samples on the two-spirals dataset for True, Full-data generator, Regression, NCDE, Frozen, EM 1 iteration and EM

Figure: Generated samples on two spirals at 60% missingness. Only co-evolving EM (right) recovers both spiral arms.

MethodImputation MSE ↓Generation (disc.) ↓
Point regression
2.22
0.390
Point imputer (NCDE)
2.46
0.252
EM, 1 iteration
3.65
0.222
Frozen stochastic
4.31
0.154
Co-evolving EM (Ours)
4.26
0.049

Table: Point-estimate metrics reward averaging; the downstream generation score reverses the ranking. The distributional Energy Score agrees with generation quality (Ours: 0.79, best). The same ordering holds in 21/21 configurations of a 7-distribution × 3-missing-rate sweep.

Method

Framework overview: STL initialization, Monte-Carlo EM loop with PSLR E-step and diffusion M-step, and the PSLR denoise-and-correct step

Figure 3: We initialize a generator from a simple decomposition of the irregular series, then refine it with a co-evolving Monte Carlo EM loop that alternates between posterior sampling (E-step) and generator updates (M-step). PSLR enables conditioning in the lifted space by iteratively denoising and correcting samples to match the observations.

Co-Evolving Monte Carlo EM

Each irregular series yts is a partial view of an unobserved complete series. Maximizing the data likelihood directly is intractable, so we alternate two steps on a diffusion generator that operates on lifted 2D (image) representations ximg = T(xts):

E-step: sample, don't average

ximg(i,k) ~ pθk−1(ximg | yts(i), Ats(i))

Draw completions from the posterior of the current generator, so they stay consistent with what the model has learned so far.

M-step: plain training

θk = argminθ 𝔼 [ λ(σ) ‖ dθ(ximg(i,k) + σε, σ) − ximg(i,k) ‖² ]

Identical to clean-data ImagenTime/EDM training: no masked loss, no auxiliary imputation head, no separate conditional model.

Posterior Sampling in Two Spaces

The prior lives in image space, but observations constrain the decoded time series. The correct observation model is therefore the composed operator

Gi = Ats(i) ∘ T−1,    yts(i) = Gi ximg

Existing diffusion posterior samplers (DPS, ΠGDM, DiffPIR, TMPD, MMPS) assume a single space. Used as-is, they correct with the wrong operator, average over modes in ambiguous regions, and drift off the manifold of valid lifted images. PSLR (Posterior Sampling for Lifted Representations) is a minimal, sampler-agnostic recipe that enforces three conditions:

C1

Operator-consistent conditioning

Enforce observations through the composed operator G = Ats ∘ T⁻¹ (mask after decoding the image back to a series), not an image-space mask approximation, so updates follow the true forward model.

C2

Uncertainty preservation

Keep the full Tweedie covariance instead of a point estimate or heuristic diagonal, with an adaptive observation-noise scale σy = c·σt, so ambiguous regions stay multimodal instead of being averaged.

C3

Manifold consistency

Project every sample back onto Range(T) via Π(x) = T(T⁻¹(x)), removing off-manifold components that would otherwise accumulate across EM iterations.

We instantiate PSLR on MMPS (full Jacobian Tweedie covariance via conjugate gradient), giving L-MMPS; applying the same recipe to TMPD gives L-TMPD, which transfers without retuning.

Full Algorithm

  1. 1

    Init

    Complete each irregular series with a cheap STL decomposition, lift it to an image, and train an initial prior pθ0 with the standard EDM objective.

  2. 2

    E-stepfor k = 1 … K

    With θk−1 frozen, sample complete lifted series ximg ~ pθ(ximg | yts) using PSLR with G = Ats ∘ T⁻¹.

  3. 3

    Projectfor k = 1 … K

    Project onto valid images, ximg ← T(T⁻¹(ximg)), and enforce the observed entries.

  4. 4

    M-stepfor k = 1 … K

    Continue training θk from the previous weights (40 epochs) on the refreshed completions, using plain clean-data EDM training.

  5. 5

    Inference

    Sample ximg ~ pθK unconditionally with standard EDM and return T⁻¹(ximg).

Modular by Design

Lift T
delay embedding (default), STFT
Operator Ats
random masks, blocks, per-cell, continuous-time
E-step sampler
any sampler meeting C1–C3 (MMPS, TMPD)
Backbone
any differentiable denoiser (EDM here)

Results

Four Corruption Regimes

All prior work evaluates a single corruption pattern. Real sensors fail for continuous periods, channels fail asynchronously, and sampling clocks are rarely uniform, so we add three new benchmarks alongside the standard one. Models are trained only on corrupted observations and evaluated by comparing unconditional samples to the fully observed data, against TimeGAN-Δt, GT-GAN, KoVAE, and ImagenI2R.

Four corruption regimes illustrated on three-channel signals

Figure 1: Four corruption regimes, all covered by a single observation operator Ats. Filled/hollow markers denote observed/missing entries.

Timestamp-level

The standard benchmark: all channels drop together at 30/50/70% of timesteps. 10 datasets × lengths 24, 96, 768.

Block missing

Contiguous gaps from sensor outages, at the start, a random location, or the end of a window.

Per-cell

Each sensor fails independently, so channels are missing asynchronously.

Continuous-time

Each channel is sampled at irregular real-valued times, off the regular grid.

Averaged Results Across Benchmarks

SettingTimeGAN-ΔtGT-GANKoVAEImagenI2ROurs
Timestep (Standard)0.4960.3870.2280.0670.055
↓ 18%
Block0.4990.4320.3210.3660.097
↓ 70%
Per-cell0.4940.4470.3410.3040.116
↓ 62%
Continuous0.4980.4270.3260.2740.043
↓ 84%

Table 1: Discriminative Score averaged across datasets for each corruption regime (lower is better). The percentage under our score is the relative improvement over the strongest baseline in each row. Our method is best in every setting and metric; prior methods, built around open-loop imputation, break down on the three new benchmarks.

Per-Dataset Results Explorer

Step 1: Choose Corruption Regime

A fraction of timesteps is dropped; all features at a dropped timestep go missing together.

Step 2: Choose Missing Rate

Step 3: Choose Metric to View

Currently viewing: Standard (timestamp) · Length 24 · 50% missing · Discriminative Score

ModelETTh1ETTh2ETTm1ETTm2WeatherElectricityEnergySineStockMujoco
TimeGAN-Δt0.4990.4990.4990.4990.4990.4980.4790.4960.4870.483
GT-GAN0.4620.3710.4070.3760.4960.3910.3170.3720.2650.270
KoVAE0.1880.0860.0570.0770.4980.4990.2980.0300.0920.117
ImagenI2R0.0320.0050.0130.0130.0350.3600.0650.0140.0070.007
Ours0.0110.0070.0190.0160.0200.3500.0420.0150.0090.019
Clean-data oracle0.0060.0060.0020.0030.0030.0050.0090.0070.0070.022

Lower is better for all metrics; the best model per dataset is highlighted. Scroll the table sideways to see every dataset. Values for our method are the minimum across the EM training history. The clean-data oracle is trained on fully observed data and serves as a reference.

Qualitative Evaluation

t-SNE embeddings and density plots for Weather, Energy and Stocks at 70% missing rate, comparing real data, ours and ImagenI2R

Figure: 2D t-SNE embeddings (top) and probability densities (bottom) at 70% missingness on Weather, Energy and Stocks. Our samples (orange) overlap the real data (blue) almost perfectly, while ImagenI2R (green) is more concentrated and slightly biased. Ours also wins every cell of the Wasserstein-distance comparison at 30%, 50% and 70%.

Ablation Study

A stochastic imputer (CSDI) halves the error of a deterministic one, and a single round of posterior sampling with extended training performs about the same. Neither is enough. Only iterating closes the gap, and a naive single-space sampler in the loop is not enough either: every PSLR condition contributes.

MethodEnergyWeatherStockAvg.
Open loop: does stochasticity or more training suffice?
Point imputer (NCDE) + ImagenTime
open loopdeterministicPSLR ✗
0.2250.2810.1490.218
CSDI + ImagenTime
open loopstochasticPSLR ✗
0.1200.0950.0700.095
Ours (1 iter, extended training)
open loopposteriorPSLR ✓
0.1400.1200.0850.115
Closed loop: does each PSLR condition matter?
Lifted-mask MMPS-EM
closed loopposteriorPSLR ✗
0.0950.0980.0960.087
w/o C1
closed loopposteriorPSLR partial
0.0650.0350.0140.038
w/o C2
closed loopposteriorPSLR partial
0.0720.0480.0220.047
w/o C3
closed loopposteriorPSLR partial
0.0780.0610.0310.056
Full method
Ours (Full PSLR-EM)
closed loopposteriorPSLR ✓
0.0510.0220.0070.027
Clean-data oracle
0.0090.0030.0070.006

Table 2: Average discriminative score (lower is better) across Energy, Weather and Stock at 30%, 50% and 70% drop rates.

Fast Convergence

EM sounds expensive, but in practice it converges in about 5 iterations: the STL warm start begins close to the data manifold, and each M-step continues from the previous weights for only 40 epochs. Time-to-best is 58% lower than KoVAE and within 2.2× of ImagenI2R.

Running-minimum discriminative score per EM iteration on Energy, ETTh1, Mujoco, Weather and Stock at 70% missing rate

Figure: Running-minimum discriminative score per EM iteration at 70% missingness (the hardest setting). Curves drop sharply in the first few iterations and plateau within 5–13 iterations.

Wall-clock time to best discriminative score (hours, RTX 3090)

30% missing2.9 EM iterations on average
KoVAE
9.89
Ours
3.34
ImagenI2R
1.48
50% missing6.6 EM iterations on average
KoVAE
7.86
Ours
3.76
ImagenI2R
1.91
70% missing5.2 EM iterations on average
KoVAE
8.25
Ours
3.68
ImagenI2R
1.57
Average4.9 EM iterations on average
KoVAE
8.67
Ours
3.59
ImagenI2R
1.65

Table: Averaged across ten datasets at length 24, including all EM iterations needed to reach the best score.

Cite Us

BibTeX Citation

@inproceedings{fadlon2026closingtheloop,
  title={Closing the Loop: Co-Evolving EM for Irregular Time Series Generation in Lifted Representations},
  author={Gal Fadlon and Idan Arbiv and Omri Azencot},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}

Quick Links

PapersoonarXivsoonCodesoon