IntentGaze: Task-Aligned Gaze Correction for Stable and Responsive Gaze-Contingent XR Article

Xuning Hu, Xinan Yan, Yichuan Zhang, Yushi Wei, Yue Li, Wolfgang Stuerzlinger, Hai-Ning Liang

Abstract:

Gaze-contingent XR systems require gaze input that is accurate, responsive, and stable under realistic head, body, and environmental motion. However, raw gaze is affected by fixation jitter, pursuit lag, saccadic transitions, and motion-induced disturbances. Filtering suppresses jitter but often trades responsiveness for stability, while gaze prediction typically estimates future measured samples for latency compensation rather than the task-aligned gaze direction needed for interaction. We formulate online XR gaze correction as causal task-aligned gaze estimation, where intended gaze is treated as the task-defined direction that users are instructed to fixate on or track. To support this formulation, we designed controlled fixation, pursuit, and saccade-related tasks and collected a multi-context gaze dataset from 35 participants, comprising 6.69 million synchronized frames of measured gaze, task-defined target directions, head/body motion signals, and gaze-behavior annotations across sitting, standing, lying down, walking, and simulated vehicle movement. Based on this dataset, we introduce IntentGaze, a motion-FiLM residual correction model that predicts current-frame task-aligned gaze corrections from recent gaze and motion histories. The model uses gaze dynamics for residual correction, motion features for contextual modulation, and saccade-aware control to preserve rapid gaze transitions. In subject-independent evaluation, IntentGaze achieved mean angular errors of 0.55° for fixation and 1.30° for pursuit, reducing errors by 26% and 18% compared with the strongest baseline. It also reduced gaze instability to 0.19° and 0.88°, corresponding to reductions of 26% and 9%. Additional analyses show that motion-conditioned modulation mainly benefits motion-rich conditions, saccade-aware control reduces transition lag to approximately one frame, reduced-sensor cross-device evaluation preserves performance on another XR headset, and replay-based dwell selection indicates potential target-acquisition benefits. The compact model further supports online deployment without dedicated acceleration in runtime XR systems. These results suggest that causal task-aligned gaze estimation can serve as a practical runtime correction layer for stable and responsive gaze-contingent XR.

Date of publication: Dec - 2026
Get Citation