I am a JPMorgan Chase Associate Professor of Computer Science in the Machine Learning Department at Carnegie Mellon University. I work in Artificial Intelligence at the intersection of Computer Vision, Machine Learning, Language Understanding and Robotics. Prior to joining MLD's faculty I spent three wonderful years as a post doctoral researcher first at UC Berkeley working with
Jitendra Malik
and then at Google Research in Mountain View working with the video group. I completed my Ph.D. in GRASP, UPenn with
Jianbo Shi
. I did my undergraduate studies at the
National Technical University of Athens
and before that I was in
Crete.
Prospective students: If you want to join CMU as PhD student, just mention my name in your application.
Otherwise, if you would like to join our group in any other capacity, please fill this form
and then please send me a short email note without any documents.
Our team's research was awarded an
Amazon Faculty Award 2020.
to support work on object manipulation across diverse environments and viewpoints
I got a
Young Inverstigator Award
from AFOSR (the Air Force Research Laboratory) to suport work that will develop intelligent multimodal surveillance systems. Special thanks to
Chris
and
Adam
for their help on proposal preparation.
Our group studies Artificial Intelligence, and specifically Machine Learning models at the intersection of Computer Vision, Language Understanding and Robotics.
Our ultimate goal is to build machines that will autonomously and in interactions with humans and with the environment acquire and continuously improve world models, that would let them reason through consequences of their and others decisions, to surpass humans in both dexterity and creativity. Topics we currently focus on include representation learning, video understanding, 2D/3D unified vision language models, generative modeling, learning simulators from data, real2sim and sim2real robot learning, reinforcement learning, continual learning.
HIGenNTO: Scalable Humanoid Interaction Generation via Noise-Space Trajectory Optimization
Lalit Jayanti, Kashu Yamazaki, Yuto Shibata, Kotaro Amaya, Katerina Fragkiadaki
HIGenNTO generates contact-rich humanoid-scene interactions by optimizing the initial noise of a pretrained text-conditioned motion model under sparse spatial, temporal, and physical constraints. The resulting motions can be tracked in simulation and used to train visuomotor policies deployed on a Unitree G1.
arXiv 2026
TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations
Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta, Fan Zhang, Tanner Schmidt,
Jakob Engel, Katerina Fragkiadaki, Adam W. Harley
TrackEverything enables dense, long-horizon 3D tracking by issuing new
tracks based on new scene geometry rather than new video frames, using
persistent 3D scene representations and geometry-based de-duplication.
arXiv 2026 project page /
paper
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
Yuxuan Kuang, Sungjae Park, Katerina Fragkiadaki, Shubham Tulsiani
Dex4D learns a task-agnostic 3D point-track-conditioned policy in simulation
that zero-shot transfers to diverse real-world dexterous manipulation tasks,
using generated videos and 4D point tracks as high-level plans.
CoRL 2026 project page /
paper
GeomVLA: Unifying Scene, Motion, and Action in 3D
Ziyin Xiong, Nikolaos Gkanatsios, Moritz Reuss, Katerina Fragkiadaki
GeomVLA unifies scene representation, latent future-motion reasoning, and robot action generation in a shared robot-centric 3D coordinate frame, achieving state-of-the-art performance on CALVIN and strong results across simulation and real-world manipulation.
CoRL 2026 project page /
paper
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
Lucy Lin*, Ayush Jain*, Yifan Liu, Katerina Fragkiadaki
Qwen-3D performs attention directly in 3D world space and grounds language in persistent 3D scene representations, unifying spatial reasoning, referential grounding, instance segmentation, and visual question answering across images and videos.
ECCV 2026 project page /
GitHub /
models
Video Diffusion Alignment via Reward Gradients
Mihir Prabhudesai, Russell Mendonca, Zheyang Qin, Katerina Fragkiadaki, Deepak Pathak
VADER aligns video diffusion models using end-to-end reward gradient backpropagation from off-the-shelf differentiable reward functions.
arxiv webpage
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
Gabriel Sarch, Lawrence Jang, Michael Tarr, William Cohen, Kenneth Marino, Katerina Fragkiadaki
A technique that enables VLM agents to take initially suboptimal demonstrations and iteratively improve them, ultimately generating high-quality trajectory data that includes both optimized actions and detailed reasoning annotations suitable for more effective in-context learning and fine-tuning. NeurIPS 2024 spotlight webpage
ODIN: A Single Model for 2D and 3D Perception
Ayush Jain, Pushkal Katara, Nikolaos Gkanatsios, Adam W. Harley, Gabriel Sarch, Kriti Aggarwal, Vishrav Chaudhary, Katerina Fragkiadaki
ODIN processes both RGB images and sequences of posed RGB-D images by alternating between 2D and 3D fusion layers using projection and unprojection from camera info. New SOTA in Scannet200.
CVPR 2024 spotlight webpage
Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following
Brian Yang, Huangyuan Su, Nikolaos Gkanatsios, Tsung-Wei Ke, Ayush Jain, Jeff Schneider, Katerina Fragkiadaki
Diffusion-ES combines trajectory diffusion models with evolutionary search and achieves SOTA performance in nuPLAN. We prompt LLMs to map language instructions to shaped reward functions, and optimize them with diffusion-ES, and solve the hardest driving scenarios.
CVPR 2024 webpage
Test-time Adaptation with Slot-Centric Models
Mihir Prabhudesai, Anirudh Goyal, Sujoy Paul, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi, Gaurav Aggarwal, Thomas Kipf, Deepak Pathak, Katerina Fragkiadaki
ICML 2023 webpage