Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

Tse, Tze Ho Elden; Feng, Runyang; Zheng, Linfang; Park, Jiho; Gao, Yixing; Kim, Jihie; Leonardis, Ales; Chang, Hyung Jin

Detailed Information

Cited 0 time in webofscience

Cited 0 time in scopus

Metadata Downloads

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

Full metadata record

DC Field	Value	Language
dc.contributor.author	Tse, Tze Ho Elden	-
dc.contributor.author	Feng, Runyang	-
dc.contributor.author	Zheng, Linfang	-
dc.contributor.author	Park, Jiho	-
dc.contributor.author	Gao, Yixing	-
dc.contributor.author	Kim, Jihie	-
dc.contributor.author	Leonardis, Ales	-
dc.contributor.author	Chang, Hyung Jin	-
dc.date.accessioned	2025-05-13T05:30:18Z	-
dc.date.available	2025-05-13T05:30:18Z	-
dc.date.issued	2025-04	-
dc.identifier.issn	2159-5399	-
dc.identifier.issn	2374-3468	-
dc.identifier.uri	https://scholarworks.dongguk.edu/handle/sw.dongguk/58328	-
dc.description.abstract	With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in representing object shape and movement using 3D bounding boxes. Additionally, the reliance on object templates at test time limits their generalisability to unseen objects. To address these challenges, we propose to leverage superquadrics as an alternative 3D object representation to bounding boxes and demonstrate their effectiveness on both template-free object reconstruction and action recognition tasks. Moreover, as we find that pure appearance-based methods can outperform the unified methods, the potential benefits from 3D geometric information remain unclear. Therefore, we study the compositionality of actions by considering a more challenging task where the training combinations of verbs and nouns do not overlap with the testing split. We extend H2O and FPHA datasets with compositional splits and design a novel collaborative learning framework that can explicitly reason about the geometric relations between hands and the manipulated object. Through extensive quantitative and qualitative evaluations, we demonstrate significant improvements over the state-of-the-arts in (compositional) action recognition. Copyright © 2025, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.	-
dc.format.extent	9	-
dc.language	영어	-
dc.language.iso	ENG	-
dc.publisher	Association for the Advancement of Artificial Intelligence	-
dc.title	Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics	-
dc.type	Article	-
dc.publisher.location	미국	-
dc.identifier.doi	10.1609/aaai.v39i7.32800	-
dc.identifier.scopusid	2-s2.0-105004065199	-
dc.identifier.wosid	001478153300083	-
dc.identifier.bibliographicCitation	Proceedings of the AAAI Conference on Artificial Intelligence, v.39, no.7, pp 7437 - 7445	-
dc.citation.title	Proceedings of the AAAI Conference on Artificial Intelligence	-
dc.citation.volume	39	-
dc.citation.number	7	-
dc.citation.startPage	7437	-
dc.citation.endPage	7445	-
dc.type.docType	Proceedings Paper	-
dc.description.isOpenAccess	Y	-
dc.description.journalRegisteredClass	foreign	-
dc.relation.journalResearchArea	Computer Science	-
dc.relation.journalWebOfScienceCategory	Computer Science, Artificial Intelligence	-
dc.relation.journalWebOfScienceCategory	Computer Science, Interdisciplinary Applications	-
dc.relation.journalWebOfScienceCategory	Computer Science, Theory & Methods	-
dc.subject.keywordAuthor	Action Recognition	-
dc.subject.keywordAuthor	Bounding-box	-
dc.subject.keywordAuthor	Collaborative Learning	-
dc.subject.keywordAuthor	Object Interactions	-
dc.subject.keywordAuthor	Object Movements	-
dc.subject.keywordAuthor	Object Pose	-
dc.subject.keywordAuthor	Object Reconstruction	-
dc.subject.keywordAuthor	Pose-estimation	-
dc.subject.keywordAuthor	Superquadrics	-
dc.subject.keywordAuthor	Unified Modeling	-

Files in This Item: There are no files associated with this item.

Appears in Collections: ETC > 1. Journal Articles

Show simple item record

qrcode

Related Researcher

Researcher Kim, Ji Hie photo

Kim, Ji Hie: College of Advanced Convergence Engineering (Department of Computer Science and Artificial Intelligence)

Read more

Altmetrics

Total Views & Downloads

RSS_1.0 RSS_2.0 ATOM_1.0

30, Pildong-ro 1-gil, Jung-gu, Seoul, 04620, Republic of Korea+82-2-2260-3114

Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.

Detailed Information

Related Researcher

Altmetrics

Total Views & Downloads

BROWSE