I am training a multimodal agent to call geometric tools (COLMAP, LoFTR, and retrieval) when a structure-from-motion pair is difficult. The research idea is Gabriele Berton's; this repository is my implementation, and it is still in progress.
Fig. 1MegaDepth matches saved at training steps 10, 20, and 30. Three figures committed in Ryukijano/agentic-sfm. Not a screen recording of the agent.