Weakly-Supervised Physically Unconstrained Gaze Estimation

Rakshit Kothari, Shalini De Mello, Umar Iqbal, Wonmin Byeon, Seonwook Park, Jan Kautz

2021-05-20CVPR 2021 1Domain Generalization Gaze Estimation

Abstract

A major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly available and can be much more easily annotated with frame-level activity labels. In this work, we tackle the previously unexplored problem of weakly-supervised gaze estimation from videos of human interactions. We leverage the insight that strong gaze-related geometric constraints exist when people perform the activity of "looking at each other" (LAEO). To acquire viable 3D gaze supervision from LAEO labels, we propose a training algorithm along with several novel loss functions especially designed for the task. With weak supervision from two large scale CMU-Panoptic and AVA-LAEO activity datasets, we show significant improvements in (a) the accuracy of semi-supervised gaze estimation and (b) cross-domain generalization on the state-of-the-art physically unconstrained in-the-wild Gaze360 gaze estimation benchmark. We open source our code at https://github.com/NVlabs/weakly-supervised-gaze.

Results

Task	Dataset	Metric	Value	Model
Gaze Estimation	Gaze360	Angular Error	12.94	ResNet-18

Related Papers

Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization2025-07-17 GLAD: Generalizable Tuning for Vision-Language Models2025-07-17 MoTM: Towards a Foundation Model for Time Series Imputation based on Continuous Modeling2025-07-17 InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing2025-07-16 From Physics to Foundation Models: A Review of AI-Driven Quantitative Remote Sensing Inversion2025-07-11 Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion2025-07-08 Prompt-Free Conditional Diffusion for Multi-object Image Augmentation2025-07-08 Integrated Structural Prompt Learning for Vision-Language Models2025-07-08