OmniHead

OmniHead is a unified video model for analysis of nonverbal facial behavior. Given a clip of 16 frames, it predicts, per frame: head gestures (nod, shake, tilt, turn, up/down), blinks, saccades, facial expressions, facial action units, valence/arousal, and 3D gaze direction. The model is a spatio-temporal model built on a Swin3D video encoder and a DETR-style, task-token cross-attention decoder with per-task MLP heads.

Overview

  • Training: OmniHead was trained on
    • CelebV-HQ β€” unlabelled, used for pretraining (~35k clips, ~6M frames)
    • s-Aff-Wild2 β€” facial expressions, valence/arousal, action units (~594 videos, ~168k frames)
    • Gaze360 β€” gaze direction (~150k frames)
    • CCDb-HG β€” head gestures (115 videos, ~1M frames)
    • VAT β€” head gesture, saccade, blink (606 videos, ~70k frames)
  • Backbone: OmniHead is adapted from Omnivore (Swin3D video encoder backbone)
  • Parameters: ~30 million
  • Task: Multi-task dynamic nonverbal facial behavior analysis: joint prediction of head gesture, blink, saccade, facial expression, action units, valence/arousal, and gaze direction from tracked head video clips.
  • Framework: PyTorch

License

Copyright (c) 2026 Idiap Research Institute

CC BY-NC-SA 4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International) This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. To view a copy of this license, visit http://creativecommons.org/licenses/by-nc-sa/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.

Minimal code to instantiate the model and perform inference:

import os
from omnihead import (
    read_and_standardize_video_info,
    extract_head_tracks,
    draw_head_tracks,
    extract_omnihead,
    draw_omnihead,
)

clip_path  = "data/my_video.mp4"
output_dir = "./outputs"
video_name = os.path.basename(clip_path).split(".")[0]

# Step 1: detect and track heads
extract_head_tracks(
    clip_path=clip_path,
    output_dir=output_dir,
    conf_thr=0.5,
    iou=0.5,
    filter_head_size=100,
    interpolate=True,
    smooth=True,
    smooth_window_size=5,
)

# Steps 2 & 3: extract OmniHead predictions + render visualization
tracks_dir = os.path.join(output_dir, video_name, "tracks")
tracks = [f for f in os.listdir(tracks_dir) if f.startswith("track_head_") and f.endswith(".npy")]

for head_track in tracks:
    head_track_path = os.path.join(tracks_dir, head_track)
    extract_omnihead(
        clip_path=clip_path,
        head_track_path=head_track_path,
        output_dir=tracks_dir,
        batch_size=32,
        num_workers=10,
    )

Refer to the OmniHead GitHub repository for full usage details, including the CLI and LMDB prediction-reading utilities.

Model performance

OmniHead (multi-task, MTL) results:

Within-dataset: CCDb head gesture F1_mi/ma (frame) 0.73/0.60, F1_mi/ma (event) 0.72/0.65; VAT blink F1 (frame/event) 0.53/0.62; VAT saccade F1 (frame/event) 0.65/0.67; Gaze360 3D gaze error (full) 12.50Β°; sAff-Wild2 expression F1_ma 0.30, valence/arousal CCC 0.44, action units F1_ma 0.44 (total score 1.18).

Cross-dataset: KTH head gesture F1_mi/ma (frame) 0.52/0.46, F1_mi/ma (event) 0.52/0.50; VAT head gesture F1 (event) mi/ma 0.60/0.34; ChildPlay head gesture F1 (event) mi/ma 0.40/0.31; ChildPlay blink F1 (frame/event) 0.41/0.48; ChildPlay saccade F1 (frame/event) 0.51/0.59; GFIE 3D gaze error (full) 21.44Β°.

Full comparison tables against SOTA baselines are available in the associated publication.

Model output structure

Per frame, per tracked head:

  • head_gesture (6,) β€” softmax probs: None, Nod, Shake, Tilt, Turn, Up_down
  • blink (2,) β€” softmax probs: No Blink, Blink
  • saccade (2,) β€” softmax probs: No Saccade, Saccade
  • expression (8,) β€” softmax probs: Neutral, Anger, Disgust, Fear, Happy, Sad, Surprise, Other
  • action_unit (12,) β€” per-AU sigmoid activations
  • valence_arousal (2,) β€” continuous regression, valence/arousal in [-1, 1]
  • gaze (3,) β€” unit-sphere gaze direction vector
  • feature_map (D,) β€” spatially pooled encoder features
  • token_<task> (D,) β€” task-token embeddings
  • head_bbox (4,) β€” [x1, y1, x2, y2] at anchor frame

Citation

If you use these model, please cite the following publication:

@inproceedings{vuillecard2026omnihead,
  title     = {OmniHead: A Unified Model for Dynamic Nonverbal Facial Behaviors},
  author    = {Vuillecard, Pierre and Odobez, Jean-Marc},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Finding (CVPRF)},
  year      = {2026},
}

Note

Model weights are released under CC BY-NC-SA 4.0 and may not be used for commercial purposes; derivative/fine-tuned models must be released under the same license with attribution. Source code is released separately under AGPL-3.0. Head gesture, saccade, and blink annotations for VAT and ChildPlay (in-the-wild) will be released separately in the future.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support