Skip to content
Gyanateet Dutta
Work·Academic·CV

Work

2025Vision

MSc Thesis: Surgical Video Prediction

MSc student, University of Leeds

My MSc project developed a VAE–Transformer model for surgical video prediction. In my experiments, it improved PSNR by 2.36 dB over the selected baselines and ran at 22 FPS with FP16 mixed precision. The code is available on GitHub.

FET-VAE pipeline with content and motion encoders, a ternary latent space, autoregressive rollout, and video reconstruction.
Fig. 1The FET-VAE architecture used for autoregressive surgical-video prediction. Architecture schematic. The t=20 results are the repository’s reported JIGSAWS test-set figures; the shape of the quality-decay curve is illustrative.
Scope

Measurements are the project’s own, on the JIGSAWS test set; longer-horizon drift and downstream workflow effects fall outside the study. The thesis is not deposited in White Rose eTheses.

Stack
VAE, Transformer, Video Prediction, Medical AI
Links
Code

Also

  • 2026MVA Rare Disease Hackathon 2026
  • 2026Causal-JEPA reproduction
  • 2026GOT-JEPA surgical tool tracking
  • 2025AIMS: Surgical Phase Detection
  • 2024Pothole Detection (arXiv)
GitHub· ORCID· Google Scholar· LinkedIn· Hugging Face· X