Multimodal learning & computer vision
Gengwei Zhang
张耕维/David
Postdoctoral Researcher
Department of Computer Science
University of North Carolina at Chapel Hill

About Me
I am a Postdoctoral Research Associate in the Department of Computer Science, University of North Carolina at Chapel Hill, hosted by Prof. Tianlong Chen. I obtained my Ph.D. from the Faculty of Engineering and Information Technology at the University of Technology Sydney (UTS), supervised by Prof. Yunchao Wei and Prof. Ling Chen.
I received my bachelor's degree in School of Data and Computer Science (currently School of Computer Science and Engineering) from Sun Yat-Sen University (SYSU). Before UTS, I spent a wonderful year working with Prof. Xiaodan Liang at SYSU as a research assistant.
My research focuses on multimodal learning, with particular interests in multimodal reasoning and image and video generation. I also work on continual learning and adapting pre-trained models to new tasks.
Research directions
Reasoning, generating, and learning across tasks.
Multimodal reasoning
Understanding and improving reasoning in multimodal models through reinforcement post-training.
Explore publicationsVisual generation
Generating images and videos with flexible and efficient generative models.
Explore publicationsContinual learning
Adapting pre-trained models to new tasks while retaining prior knowledge.
Explore publicationsMultimodal LearningMultimodal reasoning · Reinforcement post-training · Image and video generation
Continual LearningPre-trained model adaptation · Sequential fine-tuning · Knowledge retention
Selected publications
A selection of my work. * Equal contribution.
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
SLCA: Slow Learner with Classifier Alignment for Continual Learning on a Pre-trained Model
Continual Object Detection via Prototypical Task Correlation Guided Gating Mechanism
Mask Matching Transformer for Few-Shot Segmentation
Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search
Ada-Segment: Automated Multi-loss Adaptation for Panoptic Segmentation
Few-Shot Segmentation via Cycle-Consistent Transformer
Bidirectional Graph Reasoning Network for Panoptic Segmentation
Auto-Panoptic: Cooperative Multi-Component Architecture Search for Panoptic Segmentation
Get in touch
Let’s talk research.
For research conversations and academic inquiries.