Request access to HUG-VIS

Complete the Hugging Face request, sign the HUG-VIS Dataset Academic Use License, and email the signed agreement from the same official institutional email address. Requests are reviewed manually by the HUG-VIS dataset maintainers.

HUG-VIS is available only to approved applicants for non-commercial academic research under the HUG-VIS Dataset Academic Use License. The Hugging Face requester, responsible applicant, and agreement signatory must be the same eligible individual, and all submitted information must be accurate and complete. After submitting this form, download and sign the agreement and email it from the same official institutional email address. Personal email submissions are not accepted. Access is granted only after manual review by the HUG-VIS dataset maintainers.

Log in or Sign Up to review the conditions and access this dataset content.

HUG-VIS

A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

Fei Ma · Zebang Cheng · Minghui Li · Hongbo Xu · Yuyong Tan · Yihua Shao · Hanling Wang · Zhou Liu · Yuqing Gao · Dong Wang · Long Ma · Laizhong Cui · Nicu Sebe · Qi Tian

Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) · Shenzhen University · Institute of Automation, Chinese Academy of Sciences · Pengcheng Laboratory · Tongji University · Tsinghua University · The Chinese University of Hong Kong · University of Trento · Huawei



Project Page arXiv Coming Soon GitHub

TL;DR: HUG-VIS evaluates multimodal emotion recognition, human video generation, voice cloning, and human video matting on the same condition-aligned grid of 8,400 controlled human performances.

Overview of the HUG-VIS dataset construction, four benchmark tasks, and cross-task analysis

Dataset Access Form

Please follow this format before submitting the gated form. Many requests are rejected because the team information does not match the expected format.

Example Application

Field Example
Team Name GML-MMGroup
Team Leader Name Fei Ma
Team Leader Official Institutional Email mafei@gml.ac.cn
Team Leader Position / Title Researcher
Hugging Face Username your-hf-username
Team Members (comma-separated) Fei Ma, Zebang Cheng, Minghui Li
Organization / University / Company Guangming Laboratory
Country / Region China
  • Submit one request per team, not one request per member.
  • The request must be submitted by the team leader or main contact person who will also sign the agreement as the responsible applicant.
  • Use the same name, institution, position/title, and official institutional email in the online request and signed agreement. Include the exact Hugging Face username in the online request and email subject.
  • The responsible applicant must be a faculty member, researcher, or research staff member employed by a university or public/non-profit research institution. Students may not sign as the responsible applicant.
  • Personal email addresses are not accepted for the responsible applicant. The signed agreement must be sent from the official institutional email entered in the request.
  • Hugging Face grants gated access to the individual account that submits the request. Approval does not automatically grant access to every listed team member, and credentials or access tokens must not be shared.
  • Listed research-group members may work with the Dataset only under the responsible applicant's direct supervision, within the approved research environment, and under the same license terms.

Access requests are reviewed manually. Submitting the form does not guarantee approval.

Overview

Human-centered visual intelligence is inherently multimodal: facial expression, body motion, vocal prosody, and linguistic content jointly convey what a person expresses and how that expression is performed. Existing emotion, generation, speech, and matting benchmarks are usually built from different subjects, acquisition conditions, and annotations, making it difficult to compare capabilities or analyze their relationships.

HUG-VIS provides a shared, condition-aligned foundation for both understanding and generation. Thirty professional actors each complete the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol. Every performance is packaged with synchronized RGB video, noise-suppressed audio, an assigned Mandarin prompt, a verified transcript, and an alpha-matte sequence.

The benchmark covers four tasks under a common zero-shot protocol:

  • Multimodal Emotion Recognition from image, video, audio, text, and multimodal inputs.
  • Human Video Generation under audio-driven and vision-driven settings.
  • Voice Cloning with speech quality, speaker preservation, and perceptual evaluation.
  • Human Video Matting with spatial and temporal alpha-matte evaluation.

Dataset at a Glance

30 actors × 7 emotions × 4 actions × 10 utterances = 8,400 clips

Property Value
Actors 30 professional actors; gender-balanced; mean age approximately 20 years
Primary spoken and text language Mandarin Chinese
Emotion conditions Happy, Angry, Sad, Afraid, Disgusted, Surprised, and Neutral
Action templates 4 per emotion
Scenario-based utterances 10 per action; 40 per emotion; 280 per actor
Performance clips 8,400 total; 1,200 per emotion
Framing Controlled, seated half-body capture
Video 1920 × 1080; captured at 240 FPS and released at 30 FPS
Audio Synchronized with video and noise-suppressed
Text Assigned Mandarin prompt and verified transcript
Foreground supervision Alpha-matte sequence

Each actor performs an identical assignment inventory. This complete actor-by-assignment grid keeps the source material aligned across people and conditions, allowing downstream differences to be attributed more precisely to the actor, model, or evaluation criterion.

The emotion labels represent instructed, acted conditions. They should not be interpreted as the actors' spontaneous emotions, mental states, or clinical affect.

Dataset Composition

Each retained performance includes:

  • an RGB video clip;
  • synchronized, noise-suppressed audio;
  • the assigned Mandarin prompt;
  • a verified transcript; and
  • a sequence-level alpha matte.

The public repository does not yet document final archive names, split definitions, file extensions, checksums, or the on-disk schema. Users should not infer a file structure from the benchmark description alone.

Data Collection and Quality Assurance

HUG-VIS green-screen studio, capture equipment, instruction display, and prompt-guided recording setup

Acquisition Setup

Item Specification
RGB camera DJI Action 5 Pro with a fixed frontal mount
Background Uniform green screen for chroma-key matting
Lighting JHC-2000S LED with fixed color temperature and intensity
Framing Seated half-body with a constant subject-camera distance
Microphone DJI Mic Mini transmitter
Recording protocol Rest → prompted performance → return to rest

The actor begins at rest, delivers the assigned Mandarin prompt together with the corresponding emotion-consistent action, and returns to the resting pose. The shared temporal structure provides consistent boundaries for processing and evaluation while preserving natural variation across actors and performances.

Processing and Quality Control

Alpha mattes are produced in Adobe Premiere Pro through chroma-key initialization followed by sequence-level refinement. Refinement focuses on hair, fingers, clothing, self-occlusion, and motion-blurred regions. RGB, audio, text, and alpha assets are then converted to consistent conventions.

Eight professional volunteers review the retained clips for:

  • alignment among RGB video, audio, text, and alpha mattes;
  • audio-visual synchronization;
  • textual-description and transcript accuracy; and
  • segmentation errors caused by motion blur or self-occlusion.

Clips that fail these checks are discarded or returned for reprocessing.

Examples from the HUG-VIS actor-by-assignment grid varying actor, emotion, and action while preserving the rest-performance-rest sequence

Benchmark Tasks

All reported tasks follow a common zero-shot evaluation protocol: no HUG-VIS benchmark sample is used for training, fine-tuning, calibration, or model selection.

Task Inputs / Setting Evaluation
Multimodal Emotion Recognition Image frames, video, audio, text, and multimodal combinations Seven-class accuracy
Audio-driven Human Video Generation Recorded audio drives animation of a source portrait CSIM, Sync-C, Sync-D, and MOS
Vision-driven Human Video Generation Visual reference drives head or body animation LPIPS, CSIM, PSNR, SSIM, FID, and MOS
Voice Cloning Synthesized speech tied to a target actor and compared with reference audio UTMOS, DNSMOS, speaker similarity, and MOS
Human Video Matting RGB video to alpha matte MAD, MSE, gradient, connectivity, and dtSSD

The zero-shot protocol describes the accompanying benchmark evaluation. It does not, by itself, replace the dataset access conditions or define permissions for other research protocols.

Selected Benchmark Findings

  • The strongest reported video-audio-text emotion recognition result reaches 83.79% accuracy, while visual-only recognition remains substantially more difficult.
  • Human video generation leaders vary across generation regimes, comparison groups, and objective or perceptual criteria.
  • Voice-cloning rankings differ between reference-free quality prediction, speaker similarity, and human judgments.
  • BiRefNet ranks first among the evaluated matting systems on MAD, MSE, dtSSD, gradient, and connectivity errors.
  • Difficulty depends jointly on the instructed emotion, model or capability, and evaluation criterion rather than defining a global emotion ranking.

See the project repository for the latest result tables and cross-task analyses.

Intended Uses

Subject to approval and all applicable access terms, HUG-VIS is intended for:

  • non-commercial academic research in multimodal human understanding and generation;
  • zero-shot evaluation and comparative benchmarking;
  • research on multimodal emotion recognition;
  • research on audio-driven or vision-driven human video generation;
  • research on voice cloning and speaker preservation; and
  • research on spatially and temporally consistent human video matting.

Out-of-Scope and Prohibited Uses

Do not use HUG-VIS to:

  • redistribute, mirror, republish, sell, sublicense, or publicly host the dataset, annotations, or restricted derivatives without prior written permission;
  • use the dataset or derived files for commercial purposes without prior written permission;
  • identify, re-identify, contact, track, impersonate, or harm recorded participants;
  • create deceptive deepfakes, defamatory or discriminatory content, sexual content, surveillance systems, or unlawful applications;
  • share Hugging Face credentials, tokens, or another user's approved access; or
  • publicly release raw or identifiable samples or outputs that reproduce a participant's recognizable face or voice without prior written permission.

Personal and Sensitive Information

HUG-VIS contains identifiable recordings of real human participants, including faces, voices, body movements, and performed expressions. These modalities create risks involving identity inference, impersonation, unauthorized synthesis, and privacy harm.

All actors provided written informed consent before recording for research capture and authorized use of their identifiable likeness, voice, and performed behavior. Public examples released by the project are limited to uses covered by those authorizations; this does not grant dataset recipients permission to republish identifiable material.

Biases and Limitations

  • HUG-VIS intentionally prioritizes condition alignment over environmental diversity. Results may not generalize to in-the-wild backgrounds, camera viewpoints, lighting, or recording equipment.
  • Capture is fixed-front, green-screen, seated, half-body, and controlled; full-body and unconstrained motion are not represented.
  • The actors have a mean age of approximately 20 years. Beyond the reported gender balance, detailed demographic coverage is not provided, limiting population-level or fairness claims.
  • The emotion conditions are instructed and acted rather than spontaneous.
  • Spoken content, assigned prompts, and transcripts are in Mandarin Chinese; cross-lingual generalization should not be assumed.
  • Alpha mattes originate from green-screen chroma key followed by sequence-level refinement rather than an independent physical sensor.
  • Quality assurance checks alignment, synchronization, transcript accuracy, and segmentation quality, but no independent emotion-perception validity study is reported.

Request Access and Download

  1. Read the license. Download and review the HUG-VIS Dataset Academic Use License. Confirm that the proposed work and institution satisfy its non-commercial academic-use terms.
  2. Submit the Hugging Face request. Sign in to the individual Hugging Face account that should receive access and complete every field in the gated form. Provide the exact Hugging Face username that should receive access.
  3. Complete and sign the agreement. Enter the same name, institution, position/title, and official institutional email used in the online request. The responsible applicant, signatory, and Hugging Face requester must be the same eligible individual. Students may not sign as the responsible applicant.
  4. Email the signed agreement. Send it from the official institutional email entered in both forms to Zebang Cheng (zebang.cheng@gmail.com) and cc Fei Ma (mafei@gml.ac.cn). Use the subject [HUG-VIS Access] Full Name | Institution | HF username. Email the signed agreement.
  5. Wait for review. The HUG-VIS dataset maintainers match the Hugging Face request to the signed agreement and institutional-email submission. If approved, access is granted only to the Hugging Face username named in the application.

After approval, authenticate with the same individual Hugging Face account:

hf auth login
hf download GML-MMGroup/HUG-VIS \
  --repo-type dataset \
  --local-dir HUG-VIS

Inaccurate or incomplete information may lead to rejection. Violations of the approved scope or access conditions may lead to suspension or termination of access.

License and Access Terms

HUG-VIS is not released under a Creative Commons license. It is distributed under the custom HUG-VIS Dataset Academic Use License. The following summary does not replace the signed agreement:

  • The Dataset and derived models or materials may be used only for scientific, educational, and non-commercial academic purposes.
  • The Dataset, access credentials, annotations, and restricted derivatives must not be sold, transferred, sublicensed, published, uploaded, or otherwise provided to third parties.
  • The original video and audio must not be edited, manipulated, composited, dubbed, replaced, or republished as modified source data. Technical processing required for research is permitted only inside the approved research environment; processed copies and new annotations remain restricted.
  • Applicants must not identify, re-identify, contact, track, impersonate, or harm recorded participants, or use HUG-VIS for surveillance, deceptive deepfakes, defamation, discrimination, sexual content, or unlawful purposes.
  • The Dataset must be stored securely and deleted when the research ends, access is withdrawn, or deletion is requested. Access may be suspended or terminated for any breach.

Citation

The official BibTeX entry and arXiv link will be added when the paper record becomes available.

If you use HUG-VIS, please cite the official paper once the citation is released.

Contact

For dataset access and project questions:

Downloads last month
94