Abstract
Existing gaze target estimation methods fail to adequately consider depth information, primarily focusing on 2D image features while neglecting the inherent 3D spatial context that could enhance global context modeling in classroom environments. To address this limitation, we propose a depth-aware gaze target estimation framework specifically designed for classroom scenarios. Our approach consists of three key components: First, a depth estimation module is developed to handle feature information degradation. Second, we design a dual-view depth transformation method to project students’ gaze cones onto the target frame. Third, we introduce a context-aware pyramid feature extraction (CPFE) module that generates multiscale high-level feature representations to strengthen global context modeling. We also construct two datasets (MPMOCS and DVSEG) for our tasks. Experimental results on these datasets demonstrate that our method achieves significant improvements in both single-view and dual-view gaze target estimation tasks.
| Original language | English |
|---|---|
| Article number | 104533 |
| Journal | Computer Vision and Image Understanding |
| Volume | 262 |
| DOIs | |
| State | Published - Dec 2025 |
Keywords
- Classroom
- Depth information
- Dual-view
- Gaze target estimation
- Spatial transformation
Fingerprint
Dive into the research topics of 'Student gaze target estimation based on depth transformation on dual-view classroom images'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver