Skip to content
Gyanateet Dutta
Work·Academic·CV

Work

Vision

Gemma-Le: VLA Policy

I implemented a compact vision-language-action policy in a LeRobot fork, using SigLIP for vision, Gemma 3 with LoRA for language, and ScaleDP for action generation. The checkpoints were trained on the robot_sim.PickNPlace simulation dataset in LeRobot format.

Gemma-Le architecture with an action trajectory being denoised over 50 diffusion steps.
Fig. 1The signal path from multimodal inputs to the ScaleDP action head. The 50-step path is drawn, not a logged rollout. Architecture diagram. The moving trajectory is drawn, not a capture.
Scope

Evaluated on simulated pick-and-place data.

Stack
Robotics, VLA, Gemma 3, Diffusion Policy
Links
Hugging Face

Also

  • 2026MVA Rare Disease Hackathon 2026
  • 2026Causal-JEPA reproduction
  • 2026GOT-JEPA surgical tool tracking
  • 2025AIMS: Surgical Phase Detection
  • 2025MSc Thesis: Surgical Video Prediction
GitHub· ORCID· Google Scholar· LinkedIn· Hugging Face· X