Martin Q. Ma
About MeMartin Q. Ma is a Ph.D. student in the School of Computer Science at Carnegie Mellon University and previously a visiting PhD at Meta Fundamental AI Research (FAIR), advised by Louis-Philippe Morency and Ruslan Salakhutdinov. His current research focuses on post-training and inference-time scaling for video reasoning with vision-language models. Specifically, his topics include active visual perception in video chain-of-thought, where models learn to retrieve or generate video frames as tools while reasoning; efficient inference-time long-form video understanding via active frame selection; and test-time scaling with best-of-N verification for video reasoning. He also worked on understanding self-supervised learning with theoretical frameworks, and designing self-supervised algorithms for multimodal tasks. His works have been published in NeurIPS, ICLR, CVPR, ICCV, and EMNLP, and were awarded a highlight in CVPR and an oral in NeurIPS Science meets Engineering of Deep Learning workshop. |