You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning

🌐 Project • 🤗 Captioner • 🕵️ OmniFysics-Agent • 🧰 Tool: PPM
📊 OPC benchmark & Dataset • 📄 Paper • 📚 Citation


Introduction

Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. Existing captioners provide detailed visual or event-level description, but do not always preserve the physical evidence behind object interactions and state changes together.

OmniFysics-Captioner is an end-to-end, tool-free omni-modal Captioner that reads raw audio and video in a single forward pass and directly generates physics-aware audiovisual captions. Trained with the active-perception agent OmniFysics-Agent, it amortizes evidence acquisition and organization without requiring external tools at inference. Within the Agent, a physical perception model (PPM) fine-tuned on approximately 2M image-level samples serves as a dedicated tool for extracting object-interaction and state-change cues.

Qualitative physical-perception case study

Contents

Model

OmniFysics-Captioner

OmniFysics-Captioner is an end-to-end, tool-free omni-modal Captioner fine-tuned from Qwen3-Omni on Daily-Physics 50K. It reads raw audio and video in a single forward pass and directly generates physics-aware, detailed captions, amortizing the evidence-acquisition and organization capability of OmniFysics-Agent without requiring external tools at inference.

At deployment, it generates a caption from raw audio and video in one forward pass without external tools.

🤗 Fysics-AI/OmniFysics-Captioner

Performance

On VDC Detailed, a detailed video-captioning benchmark, OmniFysics-Captioner achieves 57.9% accuracy and outperforms every method compared in the paper; it also leads the compared open-source methods on DREAM-1K and Omni-Cloze.

Detailed captioning benchmark results

More importantly, captions must support downstream understanding. We have a separate question-answering model read only the generated captions and then answer questions about the original videos. Across three such evaluations, the Captioner leads all compared open-source methods; on Daily-Omni it even surpasses the strongest baseline, Gemini 3.1 Pro, showing that the preserved details translate into effective evidence for question answering.

Caption-to-QA cascade results

OmniFysics-Agent

OmniFysics-Agent forms a global event timeline from a low-cost audiovisual proxy, locates local intervals that require further inspection, and dynamically orchestrates modality-specific tools. Each observation batch is written to Evidence Memory and drives the next Plan–Execute–Observe–Reflect round, progressively refining modality choice, temporal scope, and question focus. The acquired evidence is spatiotemporally aligned and traceable, and a finalizer organizes it into a temporally coherent caption.

OmniFysics-Agent pipeline

Tool: Physical Perception Model (PPM)

Within OmniFysics-Agent, the PPM focuses on objects and physical phenomena in representative frames, perceiving physical cues such as material, contact, and deformation, and analyzing object interactions, state changes, and their potential outcomes. It is fine-tuned on approximately 2M image-level physical-perception samples to output object-centric, physical-aware evidence that complements the audio and visual tools.

The released PPM checkpoint is available on Hugging Face:

🤗 Fysics-AI/OmniFysics-Captioner/PPM

Paper

Paper and supplementary material: arXiv:2609.31714.

Citation

@article{qiu2026omnifysicscaptioner,
  title   = {OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning},
  author  = {Qiu, Kaixiang and Han, Minghao and Liu, Keliang and Liu, Yizhou and Han, Jinghang and Jiang, Yue and Wang, Shunli and Zhang, Lihua and Yang, Dingkang},
  journal = {arXiv preprint},
  year    = {2026}
}

Related Resources

Together with the Captioner, we also release OmniPhysCap (OPC), an audiovisual caption benchmark with 1,000 clips and 8,000 questions, along with a 1,000-video open subset of Daily-Physics 50K. Please refer to the dataset and benchmark card for the data and evaluation details:

📊 Fysics-AI/OmniPhysics-Caption_benchmark

License

The content of this repository is released under the Apache License 2.0 with an additional non-commercial restriction: it may be used, reproduced, and distributed for research and educational purposes only. Any commercial use is prohibited without prior written permission from the maintainers. Source videos remain subject to the licenses of their original datasets.

Downloads last month
239
Safetensors
Model size
32B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for Fysics-AI/OmniFysics-Captioner