1 briefs tagged #generative models.
The arXiv report describes audio, image and video tokenizers for text-conditioned generation.