 |
NEW!
Spatially
Grounded Long-Horizon Task Planning in the Wild
Sehun Jung*,
HyunJee Song*, Dong-Hee Kim,
Reuben Tan, Jianfeng Gao, Yong Jae
Lee,
and Donghyun Kim
(*equal
contribution)
Proceedings
of the
International Conference on Intelligent
Robots & Systems (IROS),
2026
[project page]
[arXiv]
[code]
|
|
NEW!
DocHop:
Benchmarking Out-of-domain Multi-hop Reasoning
in Information-Dense Documents
Zhuoran
Yu, Le Thien Phuc Nguyen, Jaden Park,
Xinyi Gu, Zexue He, Soochahn Lee,
Rogerio Feris, and
Yong Jae Lee
Proceedings of the International Conference
on Machine Learning (ICML),
2026
[project page] [arXiv]
[code]
|
|
NEW!
Large
Language Model Teaches Visual Students:
Cross-Modality Transfer of Fine-Grained
Conceptual Knowledge
Thomas
Liang*, Zhuoran Yu*,
and Yong Jae Lee
(*equal
contribution)
Proceedings
of the
International Conference on Machine
Learning (ICML),
2026
[arXiv]
[code]
|
 |
NEW!
MAOAM:
Unified Object
& Material
Selection with
Vision-Language
Models
Jaden
Park, Valentin
Deschaintre,
Jason Kuen,
Kangning Liu,
Iliyan
Georgiev,
Krishna Kumar
Singh, Yong
Jae Lee,
and Michael
Fischer
ACM
Transactions
on Graphics
(Proceedings
of SIGGRAPH),
2026
[project
page] [arXiv] [code]
|
 |
NEW!
Agentic
Very Long
Video
Understanding
Aniket
Rege, Arka
Sadhu, Yuliang
Li, Kejie Li,
Ramya Korlakai
Vinayak,
Yuning Chai, Yong
Jae Lee,
and Hyo Jin
Kim
Proceedings
of the Annual
Meeting of the
Association
for
Computational
Linguistics (ACL),
2026
[project
page] [arXiv] [code]
|
 |
VisualToolAgent
(VisTA): A
Reinforcement
Learning
Framework for
Visual Tool
Selection
Zeyi
Huang, Yuyang
Ji, Anirudh
Sundara Rajan,
Zefan Cai, Wen
Xiao, Haohan
Wang, Junjie
Hu, and
Yong Jae Lee
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2026
[project
page] [arXiv] [code]
|
 |
Relational
Visual
Similarity
Thao
Nguyen,
Sicheng Mo,
Krishna Kumar
Singh, Yilin
Wang, Jing
Shi, Nicholas
Kolkin, Eli
Shechtman, Yong
Jae Lee*,
and Yuheng Li*
(*equal
advising)
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2026
Best
Paper at CVPR
2026 Workshop
on Cognitive
Foundations
for Multimodal
Models
[project
page] [arXiv] [code]
|
 |
Low-Resolution
Editing is All
You Need for
High-Resolution
Editing
Junsung
Lee, Hyunsoo
Lee, Yong
Jae Lee,
and Bohyung
Han
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2026
[arXiv]
|
 |
Group
Diffusion:
Enhancing
Image
Generation by
Unlocking
Cross-Sample
Collaboration
Sicheng
Mo, Thao
Nguyen,
Richard Zhang,
Nicholas
Kolkin,
Siddharth
Srinivasan
Iyer, Eli
Shechtman,
Krishna Kumar
Singh, Yong
Jae Lee,
Bolei Zhou,
and Yuheng Li
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2026
[project
page] [arXiv] [code]
|
 |
See,
Hear, and
Understand:
Benchmarking
Audiovisual
Human Speech
Understanding
in Multimodal
Large Language
Models
Le
Thien Phuc
Nguyen*,
Zhuoran Yu*,
Samuel Low Yu
Hang, Subin
An, Jeongik
Lee, Yohan
Ban, SeungEun
Chung,
Thanh-Huy
Nguyen, JuWan
Maeng,
Soochahn Lee,
and Yong
Jae Lee
(*equal
contribution)
Findings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR
Findings),
2026
[project
page] [arXiv] [code]
|
 |
Contamination
Detection for VLMs using Multi-Modal Semantic
Perturbation
Jaden Park, Mu
Cai, Feng Yao, Jingbo Shang, Soochahn Lee, and Yong
Jae Lee
International
Conference on
Learning
Representations
(ICLR),
2026
[project
page] [arXiv] [code]
|
 |
UniTalk:
Towards Universal Active Speaker Detection in
Real World Scenarios
Le Thien Phuc Nguyen*,
Zhuoran Yu*, Khoa Quang Nhat Cao, Yuwei Guo,
Tu Ho Manh Pham, Tuan Tai Nguyen, Toan Ngo
Duc Vo, Lucas Poon, Soochan Lee, and Yong
Jae Lee
(*equal
contribution)
arXiv
2025
[project
page] [arXiv] [code]
[data]
|
 |
Learning
Compositionality from Multifaceted Synthetic
Data for Language-based Object Detection
Kwanyong Park, Sojung An, Yong
Jae Lee, and
Donghyun Kim
International
Journal of
Computer
Vision
(IJCV),
2025
[pdf]
|
 |
How
Multimodal LLMs Solve Image Tasks: A Lens on
Visual Grounding, Task Reasoning, and Answer
Decoding
Zhuoran Yu and Yong Jae Lee
Conference on
Language
Modeling (COLM),
2025
[arXiv]
|
 |
X-Fusion:
Introducing New Modality to Frozen Large
Language Models
Sicheng
Mo, Thao
Nguyen, Xun
Huang,
Siddharth
Srinivasan
Iyer, Yijun
Li, Yuchen
Liu, Abhishek
Tandon, Eli
Shechtman,
Krishna Kumar
Singh, Yong
Jae Lee,
Bolei Zhou,
and Yuheng Li
Proceedings
of the
IEEE International
Conference on
Computer
Vision
(ICCV),
2025
Best
Paper at CVPR
2025
Transformers
for Vision
(T4V) Workshop
[project
page] [arXiv] [code]
|
 |
CuRe:
Cultural Gaps in the Long Tail of Text-to-Image
Models
Aniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh
Raskar, Zhuoran Yu, Aditya Kusupati*, Yong Jae Lee*,
and Ramya Korlakai Vinayak*
(*equal
advising)
Proceedings
of the
IEEE International
Conference on
Computer
Vision
(ICCV),
2025
[project
page] [arXiv] [code]
|
 |
Customizing
Domain Adapters for Domain Generalization
Yuyang Ji, Zeyi Huang, Haohan Wang, and Yong
Jae Lee
Proceedings
of the
IEEE International
Conference on
Computer
Vision
(ICCV),
2025
[pdf]
|
 |
LLaVA-PruMerge:
Adaptive Token Reduction for Efficient Large
Multimodal Models
Yuzhang Shang*, Mu Cai*, Bingxin Xu, Yong Jae
Lee^, and Yan Yan^
(*equal contribution, ^equal advising)
Proceedings
of the
IEEE International
Conference on
Computer
Vision
(ICCV),
2025
[project
page] [arXiv] [code]
|
 |
Stay-Positive:
A Case for Ignoring Real Image Features in Fake
Image Detection
Anirudh
Sundara Rajan
and Yong Jae Lee
Proceedings of the International Conference
on Machine Learning (ICML),
2025
[project page] [arXiv]
[code]
|
|
Building a Mind Palace:
Structuring Environment-Grounded Semantic Graphs
or Effective Long Video Analysis with LLMs
Zeyi Huang*, Yuyang Ji*, Xiaofang
Wang, Nikhil Mehta, Tong Xiao, Donghyun Lee, Sigmund
VanValkenburgh, Shengxin Zha, Bolin Lai, Licheng Yu,
Ning Zhang, Yong Jae Lee‡,
and Miao Liu‡
(*equal
contribution,
‡equal
advising)
Proceedings
of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2025
[project page] [arXiv] [code]
|
|
Yo’Chameleon:
Personalized Vision and Language Generation
Thao Nguyen, Krishna Kumar Singh, Jing Shi,
Trung Bui, Yong Jae Lee*, and Yuheng Li*
(*equal
advising)
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2025
[project
page] [arXiv] [code]
|
 |
Matryoshka
Multimodal Models
Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong
Jae Lee
International
Conference on
Learning
Representations
(ICLR),
2025
[project
page] [arXiv] [code]
|
 |
LLaRA:
Supercharging Robot Learning Data for
Vision-Language Policy
Xiang Li, Cristina Mata, Jongwoo Park, Kumara
Kahatapitiya, Yoo Sung Jang, Jinghuan Shang,
Kanchana Ranasinghe, Ryan Burgert, Mu Cai, Yong
Jae Lee, and Michael S. Ryoo
International
Conference on
Learning
Representations
(ICLR),
2025
[project
page] [arXiv] [code]
|
 |
On
the Effectiveness of Dataset Alignment for Fake
Image Detection
Anirudh Sundara Rajan, Utkarsh Ojha, Jedidiah
Schloesser, and Yong Jae Lee
International
Conference on
Learning
Representations
(ICLR),
2025
[project
page] [arXiv] [code]
|
 |
Diversify, Don't
Fine-Tune: Scaling Up Visual Recognition
Training with Synthetic Images
Zhuoran Yu, Chenchen Zhu, Sean Culatana,
Raghuraman Krishnamoorthi, Fanyi Xiao, and Yong Jae Lee
Transactions
on Machine
Learning
Research (TMLR),
2025
[arXiv]
|
 |
Leveraging Large Language
Models for Scalable Vector Graphics-Driven Image
Understanding
Mu Cai*, Zeyi Huang*, Yuheng Li, Haohan Wang, and Yong Jae Lee
(*equal
contribution)
Proceedings
of the
IEEE
Winter
Conference on
Applications
of Computer
Vision (WACV),
2025
[arXiv]
|
|
Cohere3D:
Exploiting Temporal Coherence for Unsupervised
Representation Learning of Vision-based
Autonomous Driving
Yichen
Xie, Hongge Chen, Gregory P. Meyer,
Yong Jae Lee, Eric Wolff,
Masayoshi Tomizuka, Wei Zhan, Yuning
Chai, and Xin Huang
IEEE
International Conference on Robotics and Automation (ICRA),
2025
[arXiv]
|
 |
LLaVA-NeXT: Improved
reasoning, OCR, and world knowledge
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li,
Yuanhan Zhang, Sheng Shen, and Yong Jae Lee
January 2024
[blog]
|
 |
LLaVA-NeXT: A Strong
Zero-shot Video Understanding Model
Yuanhan Zhang, Bo Li, Haotian Liu, Yong Jae
Lee,
Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and
Chunyuan Li
April 2024
[blog]
|
 |
Yo'LLaVA: Your
Personalized Language and Vision Assistant
Thao Nguyen, Haotian Liu,
Mu Cai, Yuheng
Li, Utkarsh Ojha,
and
Yong Jae Lee
Neural
Information
Processing
Systems (NeurIPS),
2024
[project
page] [arXiv] [code]
|
 |
What can Foundation Models’
Embeddings do?
Xueyan Zou, Linjie Li, Jianfeng Wang, Jianwei
Yang, Mingyu Ding, Junyi Wei, Zhengyuan Yang, Feng
Li, Hao Zhang, Shilong Liu, Arul Aravinthan, Yong
Jae Lee*,
and Lijuan Wang*
(*equal
advising)
Neural
Information
Processing
Systems (NeurIPS),
2024
[arXiv]
[code]
|
 |
VGBench:
Evaluating Large Language Models on Vector
Graphics Understanding and Generation
Bocheng Zou*, Mu Cai*, Jianrui Zhang, and Yong
Jae Lee
(*equal contribution)
Conference on
Empirical
Methods in
Natural
Language
Processing (EMNLP),
2024
[project
page] [arXiv] [code]
[dataset]
|
 |
MATE:
Meet At The Embedding - Connecting Images with
Long Texts
Young Kyun Jang, Junmo Kang, Yong Jae Lee, and
Donghyun Kim
Findings
of the Conference
on Empirical
Methods in
Natural
Language
Processing (EMNLP
Findings),
2024
[arXiv]
|
 |
Removing Distributional
Discrepancies in Captions Improves Image-Text
Alignment
Yuheng Li, Haotian Liu, Mu Cai, Yijun Li, Eli
Shechtman, Zhe Lin, Yong Jae Lee,
and Krishna Kumar Singh
Proceedings
of the
European
Conference
on Computer
Vision
(ECCV),
2024
[project
page] [arXiv] [code]
|
 |
CounterCurate: Enhancing
Physical and Semantic Visio-Linguistic
Compositional Reasoning via Counterfactual
Examples
Jianrui Zhang*, Mu Cai*, Tengyang Xie, and Yong Jae Lee
(*equal
contribution)
Findings of
the
Association
for
Computational
Linguistics (ACL
Findings),
2024
[project
page] [arXiv] [code]
|
 |
Cross-Modal
Self-Supervised Learning with Effective
Contrastive Units for Point Clouds
Mu Cai, Chenxu Luo, Yong Jae Lee, and
Xiaodong Yang
IEEE/RSJ
International
Conference on
Intelligent
Robots and
Systems (IROS),
2024
[arXiv]
|
 |
Improved Baselines with
Visual Instruction Tuning (LLaVA-1.5)
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2024
(Highlight, top 2.8%)
[project
page] [arXiv] [demo]
[code]
|
 |
Making Large Multimodal
Models Understand Arbitrary Visual Prompts
(ViP-LLaVA)
Mu Cai, Haotian Liu, Siva Karthik Mustikovela,
Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2024
[project
page] [arXiv] [demo]
[code] |
 |
Edit One for All:
Interactive Batch Image Editing
Thao Nguyen, Utkarsh Ojha,
Yuheng Li, Haotian Liu, and Yong Jae Lee
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2024
[project
page] [arXiv] [code]
|
 |
Computer Vision on the
Edge: Individual Cattle Identification in
Real-Time With ReadMyCow System
Moniek Smink, Haotian Liu, Dorte Dopfer, and Yong
Jae Lee
Proceedings of the IEEE
Winter Conference on Applications of Computer
Vision (WACV), 2024
[pdf]
|
 |
Investigating the
Catastrophic Forgetting in Multimodal
Large Language Models
Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai,
Qing Qu, Yong Jae Lee, and Yi Ma
Conference
on Parsimony
and Learning (CPAL),
2024
[arXiv]
|
 |
Vinoground:
Scrutinizing LMMs over Dense Temporal Reasoning
with Short Videos
Jianrui Zhang*, Mu Cai*,
and Yong Jae Lee
(*equal
contribution)
arXiv
2024
[project
page] [arXiv] [code]
[data]
[leaderboard]
|
 |
TemporalBench:
Benchmarking Fine-grained Temporal Understanding
for Multimodal Video Models
Mu Cai, Reuben Tan, Jianrui
Zhang, Bocheng Zou, Kai Zhang, Feng Yao,
Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang
Shang, Yao Dou, Jaden Park, Jianfeng Gao^, Yong Jae Lee^,
Jianwei Yang^
(^equal
advising)
arXiv
2024
[project
page] [arXiv] [code]
[data]
[leaderboard]
|
 |
Exploring the
Capabilities of a General-Purpose Robotic
Arm in Chess Gameplay
Kazuki Shin, Sankalp Yamsani, Roman Mineyev,
Hongyu Chen, Nitish Gandi, Yong Jae Lee,
and Joohyung Kim
IEEE-RAS
International
Conference on
Humanoid
Robots (Humanoids),
2023
[pdf]
[video]
|
 |
Visual Instruction Tuning
(LLaVA)
Haotian Liu*, Chunyuan Li*, Qingyang Wu, and Yong Jae Lee
(*equal
contribution)
Neural
Information
Processing
Systems (NeurIPS),
2023
(Oral presentation, top 0.5%)
[project
page] [arXiv]
[demo] [code]
|
 |
What Knowledge Gets
Distilled in Knowledge Distillation?
Utkarsh Ojha*, Yuheng Li*, Anirudh Sundara Rajan*,
Yingyu Liang, and Yong Jae Lee
(*equal
contribution)
Neural
Information
Processing
Systems (NeurIPS),
2023
[arXiv]
|
 |
Visual Instruction
Inversion: Image Editing via Image
Prompting
Thao Nguyen, Yuheng Li, Utkarsh
Ojha, and Yong Jae Lee
Neural
Information
Processing
Systems (NeurIPS),
2023
[project
page] [arXiv]
[code]
|
 |
Segment Everything
Everywhere All at Once
Xueyan Zou‡,
Jianwei Yang‡,
Hao Zhang‡, Feng
Li‡,
Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng
Gao*, and Yong Jae Lee*
(‡,*equal
contribution)
Neural
Information
Processing
Systems (NeurIPS),
2023
[arXiv]
[demo]
[code]
|
 |
A Sentence
Speaks a Thousand Images: Domain Generalization
through Distilling CLIP with Language Guidance
Zeyi Huang, Andy
Zhou, Zijian Ling, Mu Cai, Haohan Wang, and Yong Jae Lee
Proceedings of the IEEE International
Conference on Computer Vision (ICCV),
2023
[arXiv] [code]
|
 |
GLIGEN: Open-Set
Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang
Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao,
Chunyuan Li*, and Yong Jae Lee*
(*equal
advising)
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2023
[project
page] [arXiv]
[demo]
[code]
|
 |
Generalized
Decoding for Pixel, Image, and Language
Xueyan Zou*, Zi-Yi Dou*, Jianwei
Yang*, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang
Dai, Jianfeng Wang, Lu Yuan, Nanyun Peng, Lijuan
Wang, Harkirat Behl, Yong Jae Lee‡, and
Jianfeng Gao‡
(*,‡
equal contribution)
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2023
[project
page] [arXiv]
[demo]
[code]
|
 |
Towards Universal Fake
Image Detectors that Generalize Across
Generative Models
Utkarsh Ojha*, Yuheng Li*, and Yong Jae Lee
(*equal
contribution)
Proceedings
of the
IEEE Conference
on Computer
Vision and
Pattern
Recognition
(CVPR),
2023
[project
page] [arXiv] [code]
|
 |
REACT: Learning
Customized Visual Models with Retrieval-Augmented
Knowledge
Haotian Liu, Kilho Son, Jianwei
Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee*,
and Chunyuan Li*
(*equal
advising)
Proceedings of the IEEE Conference
on Computer Vision and Pattern
Recognition (CVPR),
2023 (Highlight, top 2.5%)
[project
page] [arXiv]
[code]
|

|
InPL:
Pseudo-labeling the Inliers First for Imbalanced
Semi-supervised Learning
Zhuoran Yu, Yin Li, and Yong Jae Lee
International Conference on Learning Representations (ICLR),
2023
[arXiv]
[code]
|
 |
Generate
Anything
Anywhere in
Any Scene
Yuheng li,
Haotian Liu,
Yangming Wen,
and Yong
Jae Lee
arXiv
2023
[arXiv]
|
 |
Delving
Deeper into Anti-aliasing in ConvNets
Xueyan Zou,
Fanyi Xiao, Zhiding
Yu, Yuheng Li, and Yong
Jae Lee
International Journal of Computer Vision (IJCV),
2022 (journal
extension of our BMVC 2020
conference paper)
Invited article for best
papers of BMVC 2020
[pdf] [code]
|
 |
ELEVATER: A Benchmark and
Toolkit for Evaluating Language-Augmented Visual
Models
Chunyuan Li*, Haotian Liu*, Liunian Harold Li,
Pengchuan Zhang, Jyoti Aneja, Jianwei Yang, Ping
Jin, Houdong Hu,
Zicheng Liu, Yong
Jae Lee,
and Jianfeng Gao
(*equal
contribution)
Neural
Information Processing Systems (NeurIPS),
Datasets and Benchmarks Track, 2022
[project
page] [arXiv]
[talk
video] [toolkit]
|
 |
Masked
Discrimination for Self-Supervised Learning on
Point Clouds
Haotian Liu, Mu Cai,
and Yong Jae Lee
Proceedings of the European Conference
on Computer Vision (ECCV), 2022
[arXiv] [code]
[talk
video]
|
 |
Contrastive
Learning for Diverse Disentangled Foreground
Generation
Yuheng Li, Yijun Li, Jingwan Lu,
Eli Shechtman, Yong Jae Lee, and Krishna
Kumar Singh
Proceedings of the European Conference on
Computer Vision (ECCV), 2022
[project
page] [arXiv]
|

|
GIRAFFE
HD: A High-Resolution 3D-aware Generative Model
Yang Xue, Yuheng Li, Krishna Kumar Singh, and Yong
Jae Lee
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR),
2022
[project page] [arXiv] [code]
|

|
The Two
Dimensions of Worst-case Training and the
Integrated Effect for Out-of-domain Generalization
Zeyi Huang*, Haohan Wang*, Dong Huang, Yong Jae
Lee† and Eric Xing†
(*,† equal
contribution)
Proceedings
of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR),
2022
[arXiv] [code]
|
 |
Toward
Learning Human-aligned Cross-domain Robust Models
by Countering Misaligned Features
Haohan Wang, Zeyi Huang, Hanlin
Zhang, Yong Jae Lee,
and Eric Xing
Proceedings of the Conference on Uncertainty in
Artificial Intelligence
(UAI),
2022
[arXiv]
|

|
Equine
Pain Behaviour Classification via Self-supervised
Disentangled Pose Representation
Maheen Rashid, Sofia Broome, Katrina Ask, Elin
Hernlund, Pia Haubro Andersen, Hedvig Kjellstrom,
and Yong Jae Lee
Proceedings of the IEEE
Winter Conference on Applications of Computer
Vision (WACV), 2022
[arXiv]
|
 |
PartGAN:
Weakly-supervised Part Decomposition for Image
Generation and Segmentation
Yuheng Li, Krishna Kumar Singh,
Yang Xue, and Yong
Jae Lee
Proceedings
of the British Machine Vision Conference (BMVC),
2021
[pdf]
|

|
Collaging
Class-specific GANs for Semantic Image Synthesis
Yuheng Li, Yijun Li, Jingwan Lu,
Eli Shechtman, Yong Jae Lee, and Krishna
Kumar Singh
Proceedings of the IEEE
International Conference on Computer Vision (ICCV), 2021
[arXiv]
[talk
video]
|

|
YolactEdge:
Real-time Instance Segmentation on the Edge
Haotian
Liu*,
Rafael A. Rivera-Soto*, Fanyi
Xiao, and
Yong Jae Lee
(*equal contribution)
IEEE
International Conference on Robotics and Automation (ICRA),
2021
[arXiv] [code]
[youtube]
[talk video]
[Colab
Notebook] [Colab
Notebook (TensorRT)]
|
 |
Few-shot
Image Generation via Cross-domain Correspondence
Utkarsh Ojha, Yijun Li, Jingwan Lu, Alexei A. Efros, Yong
Jae Lee, Eli Shechtman, and Richard Zhang
Proceedings
of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR),
2021
[project
page] [arXiv] [code]
|
 |
Progressive
Temporal Feature Alignment Network for Video
Inpainting
Xueyan Zou, Linjie Yang, Ding Liu, and Yong Jae
Lee
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR),
2021
[arXiv] [code]
[youtube]
|

|
Generating
Furry Cars: Disentangling Object Shape and
Appearance across Multiple Domains
Utkarsh Ojha, Krishna Kumar Singh, and Yong Jae
Lee
International Conference on Learning Representations (ICLR),
2021
[project
page] [open
review] [arXiv]
[talk video]
|

|
SinGAN-GIF:
Learning a Generative Video Model from a Single GIF
Rajat Arora and Yong Jae Lee
Proceedings of
the IEEE Winter Conference on
Applications of Computer Vision (WACV),
2021
[project
page] [pdf]
[talk
video]
|
 |
Seeing the Unseen: Predicting
the First-Person Camera Wearer's Location and Pose
in Third-Person Scenes
Yangming Wen, Krishna Kumar Singh, Markham Anderson,
Wei-Pang Jan, and Yong Jae Lee
International
Workshop on Egocentric Perception, Interaction
and Computing (EPIC), ICCV 2021
[pdf]
|

|
Elastic-InfoGAN:
Unsupervised Disentangled Representation Learning in
Class-Imbalanced Data
Utkarsh Ojha, Krishna Kumar Singh, Cho-Jui Hsieh, and
Yong Jae Lee
Neural Information Processing Systems (NeurIPS),
2020
[project
page] [arXiv]
[code] |
 |
YOLACT++:
Better Real-time Instance Segmentation
Daniel Bolya*, Chong Zhou*, Fanyi
Xiao, and
Yong Jae Lee
(*equal contribution)
IEEE
Transactions on Pattern Analysis and Machine
Intelligence (TPAMI), 2020 (journal extension
of our ICCV 2019
conference paper with improved models)
[arXiv] [code]
|
 |
Delving
Deeper into Anti-aliasing in ConvNets
Xueyan Zou,
Fanyi Xiao, Zhiding
Yu, and Yong Jae
Lee
Proceedings of the British Machine Vision Conference (BMVC),
2020 (Oral presentation)
Best Paper
Award
[project
page] [arXiv]
[code] [talk
video] |
 |
Password-conditioned
Anonymization and Deanonymization with Face Identity
Transformers
Xiuye Gu, Weixin Luo,
Michael Ryoo, and
Yong Jae Lee
Proceedings of the European Conference on
Computer Vision (ECCV), 2020
[arXiv]
[code]
[demo]
[1
min talk video] [10
min talk video]
|
 |
Boxer: Preventing Fraud by
Scanning Credit Cards
Zainul Abi Din, Hari
Venugopalan, Jaime Park, Andy Li, Weisu Yin, Haohui
Mai, Yong Jae Lee, Steven Liu, and Samuel T.
King
Proceedings of the USENIX
Security Symposium (USENIX Security), 2020
[pdf] [project
page] [talk
video] |
 |
MixNMatch:
Multifactor Disentanglement and Encoding for
Conditional Image Generation
Yuheng Li, Krishna Kumar Singh, Utkarsh Ojha, and
Yong Jae Lee
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR), 2020
[arXiv]
[code]
[youtube] [talk
video]
|
 |
Don’t Judge
an Object by Its Context: Learning to Overcome
Contextual Bias
Krishna Kumar Singh, Dhruv
Mahajan, Kristen Grauman, Yong Jae Lee,
Matt Feiszli, and Deepti
Ghadiyaram
Proceedings
of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2020 (Oral presentation)
[arXiv]
[project
page]
|

|
Instance-aware,
Context-focused, and Memory-efficient
Weakly-supervised Object Detection
Zhongzheng Ren, Zhiding Yu,
Xiaodong Yang, Ming-Yu Liu, Yong Jae Lee,
Alexander Schwing, and Jan Kautz
Proceedings
of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2020
[arXiv]
[project
page] [code]
|

|
Action
Graphs: Weakly-supervised Action Localization with
Graph Convolution Networks
Maheen Rashid, Hedvig
Kjellström, and Yong Jae Lee
Proceedings of the IEEE
Winter Conference on Applications of Computer Vision (WACV),
2020
[arXiv]
[code]
|
 |
Audiovisual
SlowFast Networks for Video Recognition
Fanyi Xiao, Yong Jae Lee,
Kristen Grauman, Jitendra Malik, and
Christoph Feichtenhofer
arXiv 2019
[arXiv]
|

|
YOLACT:
Real-time Instance Segmentation
Daniel Bolya, Chong Zhou, Fanyi
Xiao, and
Yong Jae Lee
Proceedings of the IEEE
International Conference on Computer Vision (ICCV), 2019 (Oral presentation)
Most Innovative Award, COCO
Object Detection Challenge, ICCV 2019
[arXiv] [code] [pdf]
[talk
video]
|

|
Identity from here, Pose
from there: Self-supervised Disentanglement and
Generation of Objects using Unlabeled Videos
Fanyi Xiao, Haotian Liu, and
Yong Jae Lee
Proceedings of the IEEE
International Conference on Computer Vision (ICCV), 2019
[pdf]
|

|
FineGAN:
Unsupervised Hierarchical Disentanglement for
Fine-Grained Object Generation and Discovery
Krishna Kumar Singh*, Utkarsh Ojha*, and
Yong Jae Lee
(*equal contribution)
Proceedings
of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2019 (Oral
presentation)
[project
page] [pdf] [arXiv]
[code]
[youtube]
[talk
video]
|

|
You
reap what you sow: Using Videos to Generate High
Precision Object Proposals for Weakly-supervised
Object Detection
Krishna Kumar Singh and
Yong Jae Lee
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR), 2019
[project
page] [pdf]
[code]
|

|
HPLFlowNet:
Hierarchical Permutohedral Lattice FlowNet for Scene
Flow Estimation on Large-scale Point Clouds
Xiuye Gu, Yijie Wang, Chongruo Wu, Yong Jae Lee,
and Panqu Wang
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR), 2019
[pdf] [supp]
[code]
|
 |
Video
Object Detection with an Aligned Spatial-Temporal
Memory
Fanyi
Xiao and Yong Jae Lee
Proceedings of the European Conference on
Computer Vision (ECCV),
2018
[project page] [pdf] [code]
|
 |
Learning to Anonymize Faces for Privacy
Preserving Action Detection
Zhongzheng
Ren, Yong Jae Lee, and Michael Ryoo
Proceedings of the European Conference on
Computer Vision (ECCV),
2018
[project page] [pdf] [youtube]
|

|
DOCK:
Detecting Objects by transferring
Common-sense Knowledge
Krishna Kumar Singh,
Santosh Divvala, Ali Farhadi, and Yong Jae
Lee
Proceedings of the European Conference on
Computer Vision (ECCV),
2018
[project
page] [pdf] [code]
|

|
A
Visual Attention Grounding Neural Model for
Multimodal Machine Translation
Mingyang Zhou, Runxiang
Cheng, Yong Jae
Lee, and Zhou Yu
Proceedings of the Conference on
Empirical Methods in Natural Language Processing (EMNLP), 2018 (Oral
presentation)
[pdf] |

|
Cross-Domain
Self-supervised Multi-task Feature Learning using
Synthetic Imagery
Zhongzheng
Ren and Yong Jae Lee
Proceedings of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2018
[project
page] [pdf]
[code]
|
 |
Who Will Share My
Image? Predicting the Content Diffusion Path in
Online Social Networks
Wenjian
Hu, Krishna Kumar Singh*, Fanyi Xiao*, Jinyoung Han,
Chen-Nee Chuah, and Yong Jae
Lee
(*equal contribution)
Proceedings
of the ACM International Conference on Web Search and
Data Mining (WSDM), 2018
[pdf]
|
 |
Can a Machine Learn to
See Horse Pain? An Interdisciplinary Approach
Towards Automated Decoding of Facial Expressions of
Pain in the Horse
Pia Andersen, Karina Gleerup, Jennifer Wathan, Britt
Coles, Hedvig Kjellström, Sofia Broome, Yong Jae
Lee, Maheen Rashid, Claudia Sonder, Erika
Rosenberger, and Deborah Forster
International Conference on Methods and Techniques in
Behavioral Research (Measuring Behavior), 2018
[pdf]
|
 |
What Should I Annotate?
An Automatic Tool for Finding Video Segments for
EquiFACS Annotation
Maheen Rashid, Sofia Broome, Pia Andersen, Karina
Gleerup, and Yong Jae Lee
International Conference on Methods and Techniques in
Behavioral Research (Measuring Behavior), 2018
[pdf]
|

|
Hide-and-Seek:
Forcing a Network to be Meticulous for
Weakly-supervised Object and Action Localization
Krishna
Kumar Singh and Yong Jae
Lee
Proceedings of the IEEE International Conference on
Computer Vision (ICCV), 2017
[project
page] [pdf]
[supp] [code]
|
 |
Weakly-supervised
Visual Grounding of Phrases with Linguistic
Structures
Fanyi
Xiao, Leonid Sigal, and Yong Jae
Lee
Proceedings of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2017
[project
page] [pdf]
|
 |
Interspecies
Knowledge Transfer for Facial Keypoint Detection
Maheen Rashid, Xiuye Gu, and Yong Jae
Lee
Proceedings of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2017
[project
page] [pdf] [code]
[data]
|
 |
Identifying
First-Person Camera Wearers in Third-Person Videos
Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar
Singh, Yong Jae Lee, David Crandall and
Michael Ryoo
Proceedings of the IEEE Conference on
Computer Vision and Pattern Recognition (CVPR), 2017
[pdf]
|

|
Who
Moved My Cheese? Automatic Annotation of Rodent
Behaviors with Convolutional Neural Networks
Zhongzheng Ren, Adriana Noronha, Annie Vogel Ciernia,
and Yong Jae Lee
Proceedings of the Winter Conference on Applications
of Computer Vision (WACV),
2017
[project
page] [pdf] [code]
[data]
|
|
Analyzing the Adoption
and Cascading Process of OSN-Based Gifting
Applications: An Empirical Study
M. Rezaur Rahman, Jinyoung Han, Yong Jae Lee,
and Chen-Nee Chuah
ACM Transactions on the Web (TWEB),
2017
[pdf]
|

|
End-to-End
Localization and Ranking for Relative Attributes
Krishna Kumar Singh and Yong Jae Lee
Proceedings of the European Conference on
Computer Vision (ECCV), 2016
[project
page] [pdf]
[code]
|

|
Track
and Transfer: Watching Videos to Simulate Strong
Human Supervision for Weakly-Supervised Object
Detection
Krishna Kumar Singh, Fanyi Xiao, and Yong Jae
Lee
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR),
2016
[project
page] [pdf]
[arXiv
(with more results)] [code] |
 |
Track and Segment: An Iterative Unsupervised Approach for
Video Object Proposals
Fanyi Xiao and Yong Jae Lee
Proceedings of the IEEE Conference on Computer Vision and
Pattern Recognition (CVPR),
2016 (Spotlight presentation)
[project
page] [pdf]
[code]
|

|
Localizing
and Visualizing Relative Attributes
Fanyi Xiao and Yong Jae
Lee
Springer Book Chapter on Visual Attributes, 2016
[pdf]
[code]
|

|
Discovering
Mid-level Visual Connections in Space and Time
Yong Jae Lee, Alexei
A. Efros, and Martial Hebert
Springer Book Chapter on Visual Analysis and
Geo-Localization of Large Scale Imagery, 2016
[pdf]
[code]
[data]
|

|
Discovering
the Spatial Extent of Relative Attributes
Fanyi Xiao and Yong Jae Lee
Proceedings of the IEEE International Conference on
Computer Vision (ICCV), 2015 (Oral presentation)
[project
page] [pdf]
[slides]
[code]
[video
presentation]
|

|
FlowWeb:
Joint Image Set Alignment by Weaving Consistent,
Pixel-wise Correspondences
Tinghui Zhou, Yong Jae Lee, Stella X. Yu, and
Alexei A. Efros
Proceedings of the IEEE Conference on Computer Vision
and Pattern Recognition (CVPR), 2015 (Oral presentation)
[project
page] [pdf]
[code]
|

|
Predicting
Important Objects for Egocentric Video Summarization
Yong Jae Lee and Kristen Grauman
International Journal of Computer Vision (IJCV),
2015
[project page]
[pdf]
[arXiv]
[data]
|

|
Weakly-supervised
Discovery of Visual Pattern Configurations
Hyun Oh Song, Yong Jae Lee, Stefanie Jegelka,
and Trevor Darrell
Neural Information Processing Systems (NIPS),
2014
[pdf]
|

|
AverageExplorer:
Interactive Exploration and Alignment of Visual Data
Collections
Jun-Yan Zhu, Yong Jae Lee, and Alexei A. Efros
ACM Transactions on Graphics (Proceedings of SIGGRAPH),
2014 (Oral presentation)
[project
page] [pdf]
[youtube]
[See article
in The New Yorker]
|

|
Style-aware Mid-level
Representation for Discovering Visual Connections in
Space and Time
Yong Jae Lee, Alexei A. Efros, and Martial
Hebert
Proceedings of the IEEE International Conference on
Computer Vision (ICCV), 2013 (Oral presentation)
[project page] [pdf]
[slides] [code]
[data] [video
presentation]
|

|
Discovering
Important People and Objects for Egocentric Video
Summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen
Grauman
Proceedings of the IEEE Conference on Computer Vision
and Pattern Recognition (CVPR), 2012
[project
page] [pdf]
[supp]
[extended
abstract] [data]
|

|
Object-Graphs
for Context-Aware Visual Category Discovery
Yong Jae Lee and Kristen Grauman
IEEE Transactions on Pattern Analysis and Machine
Intelligence (TPAMI), 2012
[project
page] [pdf]
[code]
|

|
Key-Segments
for Video Object Segmentation
Yong Jae Lee, Jaechul Kim, and Kristen Grauman
Proceedings of the IEEE International Conference on
Computer Vision (ICCV), 2011
[project
page] [pdf]
[code]
[data]
|

|
ShadowDraw:
Real-Time User Guidance for Freehand Drawing
Yong Jae Lee, Larry Zitnick, and Michael Cohen
ACM Transactions on Graphics (Proceedings of SIGGRAPH),
2011 (Oral presentation)
[project
page] [pdf]
[slides]
[video]
[youtube]
[data]
|

|
Face
Discovery with Social Context
Yong Jae Lee and Kristen Grauman
Proceedings of the British Machine Vision Conference (BMVC),
2011
[project
page] [pdf]
[extended
abstract]
|

|
Learning
the Easy Things First: Self-Paced Visual Category
Discovery
Yong Jae Lee and Kristen Grauman
Proceedings of the IEEE Conference on Computer Vision
and Pattern Recognition (CVPR), 2011
[project
page] [pdf]
|

|
Object-Graphs
for Context-Aware Category Discovery
Yong Jae Lee and Kristen Grauman
Proceedings of the IEEE Conference on Computer Vision
and Pattern Recognition (CVPR), 2010 (Oral presentation)
[project
page] [pdf]
[supp]
[slides]
[code]
|

|
Collect-Cut:
Segmentation with Top-Down Cues Discovered in
Multi-Object Images
Yong Jae Lee and Kristen Grauman
Proceedings of the IEEE Conference on Computer Vision
and Pattern Recognition (CVPR), 2010
[project
page] [pdf]
[supp]
[data]
|

|
Foreground
Focus: Unsupervised Learning from Partially Matching
Images
Yong Jae Lee and Kristen Grauman
International Journal of Computer Vision (IJCV),
2009
[project
page] [pdf]
|

|
Shape
Discovery from Unlabeled Image Collections
Yong Jae Lee and Kristen Grauman
Proceedings of the IEEE Conference on Computer Vision
and Pattern Recognition (CVPR), 2009
[project
page] [pdf]
[supp]
|

|
Foreground
Focus: Finding Meaningful Features in Unlabeled
Images
Yong Jae Lee and Kristen Grauman
Proceedings of the British Machine Vision Conference (BMVC),
2008 (Oral presentation)
[project
page] [pdf]
[slides]
|

|
Ray-based
Color Image Segmentation
Changhai Xu, Yong Jae Lee, and Benjamin
Kuipers
Proceedings of the Canadian Conference on Computer and
Robot Vision (CRV), 2008
[pdf]
|