Capabilities
What human-centric capabilities are required for intelligent systems operating in open environments?
Connecting perception, interaction, multimodal reasoning, embodied intelligence, open-world generalization, and reliable evaluation.
Computer vision and AI systems increasingly need to understand people whose actions, interactions, goals, and intentions unfold in complex environments. Progress spans activity understanding, human-object interaction, affordance learning, multimodal grounding, foundation models, embodied intelligence, open-world learning, and reliability evaluation.
HOUR—Human-Centric Open-World Understanding and Reasoning—brings these complementary directions together under a shared goal. The workshop will connect perception, spatial grounding, interaction and intent reasoning, generalization, and reliability diagnosis while encouraging new collaborations across communities.
What human-centric capabilities are required for intelligent systems operating in open environments?
How should visual, linguistic, gestural, and contextual evidence be combined when observations are incomplete or ambiguous?
How should benchmarks reveal when and why systems fail rather than reporting only aggregate accuracy?
Talk titles and the final program will be announced after the workshop schedule is assigned.
Research on computer vision and machine learning, focusing on video understanding, multimodal perception, embodied AI, and visual recognition.
Research on human-centric computer vision and video understanding, with a focus on contextual AI systems for understanding and assisting with human activity.
Research on multimodal learning, computer vision, and efficient learning systems, including foundation models, human-centric AI, and privacy-preserving learning.
Research on computer vision, machine learning, and biometrics, with a focus on human recognition, face-related analysis, and 3D vision.
Research on multimodal video understanding, particularly egocentric video and vision-language methods for retrieval, grounding, and captioning.
HOUR welcomes original research, benchmark and dataset studies, diagnostic analyses, position papers, and works in progress.
Submissions of at least five pages following the WACV format. Accepted archival papers will be eligible for publication in the WACV 2027 Workshop Proceedings.
Extended abstracts and position papers presenting emerging ideas, findings, resources, or community perspectives. These submissions will not appear in the proceedings.
Workshop website and call for papers live
Paper and extended-abstract submission deadline
Author notification deadline for archival papers
Accepted-paper metadata due to IEEE
Camera-ready deadline for archival papers
HOUR at WACV 2027
Deadlines are based on the WACV 2027 workshop timeline. Additional submission dates will be announced.
Exact times and presentation titles will be added after WACV assigns the workshop session.
Opening remarks and framing questions for HOUR
Perspectives spanning human-centric vision, motion, egocentric understanding, and embodied AI
Selected archival and non-archival contributions
Interactive presentation of accepted work
Workshop summary and future community activities
Works on robust human-centric AI, computer vision, and multimodal foundation models.
Works on vision-language methods for human-object interaction and affordance grounding.
Works on robust foundation-model adaptation, continual learning, and robotics.
Works on vision-language models, 3D perception, scene understanding, and robotics.
Works on machine learning, computer vision, autonomous driving, and reliable AI.
Works on human-object interaction detection, evaluation, and benchmark development.
Leads research on healthcare AI, multimodal learning, and foundation models.
Works on computer vision, deep learning, generative AI, and robust visual understanding.
We will soon invite researchers with relevant expertise to express interest in reviewing for HOUR.
For workshop and submission questions, contact the organizing team.
hour.workshops@gmail.com