My MSc project developed a VAE–Transformer model for surgical video prediction. In my experiments, it improved PSNR by 2.36 dB over the selected baselines and ran at 22 FPS with FP16 mixed precision. The code is available on GitHub.
Fig. 1The FET-VAE architecture used for autoregressive surgical-video prediction. Architecture schematic. The t=20 results are the repository’s reported JIGSAWS test-set figures; the shape of the quality-decay curve is illustrative.