Abstract
Temporal Action Segmentation (TAS) is fundamental to video understanding, aiming to recognize and segment actions within untrimmed videos. While over-segmentation errors have received considerable attention, the equally critical issue of out of-context errors stemming from the neglect of implicit causal relationships between actions remains under-explored. However, constraining these unknown causal relationships at the frame level—including identifying causal links and strengths—presents a significant challenge, leading to notable research gaps. To tackle these challenges, we define a Temporal Action Causal Model (TACM) for the TAS task as declarative guidance, which efficiently injects the causal relationship between actions into frame variables in a generated manner. To this end, we propose the Causal Action Segmentation Refiner (CaASR), a refinement framework designed to mitigate out-of-context errors in segmentation results of backbone models. By reconstructing the causal generation process of post-segmentation frame variables, CaASR constrains causal relationships between actions and ensures the reliability of generated causal relationships. Extensive experiments demonstrate that CaASR significantly enhances both segmentation performance and interpretability across various backbone models, including state-of-the-art methods.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- causal generation process
- causal inference
- causal representation learning
- Temporal action segmentation
Fingerprint
Dive into the research topics of 'CaASR: A Causal Lens for Refining Temporal Action Segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver