OmniHead
OmniHead is a unified video model for analysis of nonverbal facial behavior. Given a clip of 16 frames, it predicts, per frame: head gestures (nod, shake, tilt, turn, up/down), blinks, saccades, facial expressions, facial action units, valence/arousal, and 3D gaze direction. The model is a spatio-temporal model built on a Swin3D video encoder and a DETR-style, task-token cross-attention decoder with per-task MLP heads.
Overview
- Training: OmniHead was trained on
- CelebV-HQ β unlabelled, used for pretraining (~35k clips, ~6M frames)
- s-Aff-Wild2 β facial expressions, valence/arousal, action units (~594 videos, ~168k frames)
- Gaze360 β gaze direction (~150k frames)
- CCDb-HG β head gestures (115 videos, ~1M frames)
- VAT β head gesture, saccade, blink (606 videos, ~70k frames)
- Backbone: OmniHead is adapted from Omnivore (Swin3D video encoder backbone)
- License: CC BY-NC 4.0
- Parameters: ~30 million
- Task: Multi-task dynamic nonverbal facial behavior analysis: joint prediction of head gesture, blink, saccade, facial expression, action units, valence/arousal, and gaze direction from tracked head video clips.
- Framework: PyTorch
License
Copyright (c) 2026 Idiap Research Institute
CC BY-NC-SA 4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International) This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. To view a copy of this license, visit http://creativecommons.org/licenses/by-nc-sa/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.
Minimal code to instantiate the model and perform inference:
import os
from omnihead import (
read_and_standardize_video_info,
extract_head_tracks,
draw_head_tracks,
extract_omnihead,
draw_omnihead,
)
clip_path = "data/my_video.mp4"
output_dir = "./outputs"
video_name = os.path.basename(clip_path).split(".")[0]
# Step 1: detect and track heads
extract_head_tracks(
clip_path=clip_path,
output_dir=output_dir,
conf_thr=0.5,
iou=0.5,
filter_head_size=100,
interpolate=True,
smooth=True,
smooth_window_size=5,
)
# Steps 2 & 3: extract OmniHead predictions + render visualization
tracks_dir = os.path.join(output_dir, video_name, "tracks")
tracks = [f for f in os.listdir(tracks_dir) if f.startswith("track_head_") and f.endswith(".npy")]
for head_track in tracks:
head_track_path = os.path.join(tracks_dir, head_track)
extract_omnihead(
clip_path=clip_path,
head_track_path=head_track_path,
output_dir=tracks_dir,
batch_size=32,
num_workers=10,
)
Refer to the OmniHead GitHub repository for full usage details, including the CLI and LMDB prediction-reading utilities.
Model performance
OmniHead (multi-task, MTL) results:
Within-dataset: CCDb head gesture F1_mi/ma (frame) 0.73/0.60, F1_mi/ma (event) 0.72/0.65; VAT blink F1 (frame/event) 0.53/0.62; VAT saccade F1 (frame/event) 0.65/0.67; Gaze360 3D gaze error (full) 12.50Β°; sAff-Wild2 expression F1_ma 0.30, valence/arousal CCC 0.44, action units F1_ma 0.44 (total score 1.18).
Cross-dataset: KTH head gesture F1_mi/ma (frame) 0.52/0.46, F1_mi/ma (event) 0.52/0.50; VAT head gesture F1 (event) mi/ma 0.60/0.34; ChildPlay head gesture F1 (event) mi/ma 0.40/0.31; ChildPlay blink F1 (frame/event) 0.41/0.48; ChildPlay saccade F1 (frame/event) 0.51/0.59; GFIE 3D gaze error (full) 21.44Β°.
Full comparison tables against SOTA baselines are available in the associated publication.
Model output structure
Per frame, per tracked head:
head_gesture(6,) β softmax probs: None, Nod, Shake, Tilt, Turn, Up_downblink(2,) β softmax probs: No Blink, Blinksaccade(2,) β softmax probs: No Saccade, Saccadeexpression(8,) β softmax probs: Neutral, Anger, Disgust, Fear, Happy, Sad, Surprise, Otheraction_unit(12,) β per-AU sigmoid activationsvalence_arousal(2,) β continuous regression, valence/arousal in [-1, 1]gaze(3,) β unit-sphere gaze direction vectorfeature_map(D,) β spatially pooled encoder featurestoken_<task>(D,) β task-token embeddingshead_bbox(4,) β [x1, y1, x2, y2] at anchor frame
Citation
If you use these model, please cite the following publication:
@inproceedings{vuillecard2026omnihead,
title = {OmniHead: A Unified Model for Dynamic Nonverbal Facial Behaviors},
author = {Vuillecard, Pierre and Odobez, Jean-Marc},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Finding (CVPRF)},
year = {2026},
}
Note
Model weights are released under CC BY-NC-SA 4.0 and may not be used for commercial purposes; derivative/fine-tuned models must be released under the same license with attribution. Source code is released separately under AGPL-3.0. Head gesture, saccade, and blink annotations for VAT and ChildPlay (in-the-wild) will be released separately in the future.