Back

Explorative modeling: Train on the best of K guesses

42 points4 hoursalexiglad.github.io
ollin1 hour ago

This paper shows a nice integration of older winner-take-all ideas for learning K-modal generative models (see e.g. https://arxiv.org/abs/1612.00197, https://arxiv.org/abs/2211.14286) into modern diffusion/flow pipelines. As-implemented, I think it has some downsides:

1. K-1 extra forward passes during training

2. Inaccurate sampling behavior (will sample all K modes with equal likelihood, rather than sampling them proportionally)

However, I think both of these downsides can be mitigated by adjusting the implementation a bit more (you can have the model predict K modes jointly in one forward pass, along with probabilities of each being the min-loss mode, which you can then use for properly-weighted mode sampling at inference).

That said, I'm not entirely sure if this strategy is as generally applicable as the authors are hoping. In particular:

1. For highly-conditional image generation (like modern commercial diffusion pipelines, which use a big LLM preprocessor), most of the low-frequency color/layout decisions are already made for you by the conditioning signal. The diffusion process mostly needs to generate high-frequency details, for which there are a huge number of equally-valid modes.

2. For LLMs themselves, the sequence-generation process is already factored into a discrete classification problem, and there's no mode smearing issue to fix.

kamranjon2 hours ago

This is amazing and I think will probably end up being a pretty important development.

I was just reading this great breakdown of how diffusion Gemma works: https://newsletter.maartengrootendorst.com/p/a-visual-guide-...

In reference to the difficulties with applying this to autoregressive LLMs - I wonder if these type of hybrids might be a good candidate for this approach.

justinnk1 hour ago

I agree, this seems like an important discovery. And it hits a sweetspot of being effective and very simple to implement. Have a look at the pseudocode comparison at the bottom of the accompanying github page: https://explorative-modeling.github.io/

Also, a very well-structured and easy to follow blog post by the author. I very much enjoyed reading it!

FeepingCreature3 hours ago

If their numbers are real and hold up at scale, every image model trained before this is now obsolete.

ltsSmitty2 hours ago

I don't have enough context to know if this is actually cool or not, but it seems like it!

zimablue2253 minutes ago

This is just GRPO (proposed by DeepSeek), which similarly samples many plausible generations, selects the best of K, and trains that sample.

Minibatch OT in flow matching also has a very similar mechanism, where samples from a noise distribution are matched to the closest data point.

There is a lot of prior work here that the authors neglect to discuss, which portrays this work as more novel than it actually is.

yorwba32 minutes ago

GRPO increases the likelihood of samples that are better than average, not just the single best, and decreases that of samples that are worse than average. This method doesn't even involve an explicit likelihood, so it's a completely different mechanism.

A comparison with minibatch optimal transport is in appendix A.2 of the paper.

zimablue2223 minutes ago

You're right about the original GRPO proposal, but there are simplified variants that do just use best of K sampling.

GRPO (or GRPO like approaches) for diffusion/flow matching similarly can be likelihood free.

imenani41 minutes ago

The relation to current RLVR methods I think is interesting, they do discuss it a bit but I would be curious to see more about this as well. Quote from the paper:

Exploration beyond Pretraining. The mode collapse XMs address during pretraining also often shows up in post-training, where RL fine-tuning is known to sharpen models onto a narrow set of behaviors [76]. Recent fixes such as pass@krewards [77] and best-of-N-aware fine-tuning [78] can be seen through our lens as Forward XM, with a verifier standing in for ground truth data. These fixes act only during post-training, though; pretraining with exploration may yield base models that capture more modes in the first place, leaving RL more to select among.

zimablue2236 minutes ago

That's fair.

On the other hand, Minibatch OT (optimal transport) was one of the more fundamental advancements early on in flow matching & rectified flow models.

This best of K approach effectively discards matches that would otherwise be preserved by optimal transport assignment.

I still think the authors here are too broad regarding their claimed novelty.

cousin_it1 hour ago

No mention of GANs? As far as I remember, their images felt much more like "sampled from training data" than "averaged from training data". The idea was that you train a "discriminator" that tells generated images apart from real ones, and make the "generator" try to fool it. But somehow it lost to diffusion models, and now all AI imagery looks like slop, when GAN outputs (though imperfect) didn't look nearly as slop. I don't really know what happened though.