Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

RGB-D egocentric segmentation of human bodies for XR applications

View through CrossRef
Introduction Video-based self-avatars represent a promising approach for displaying users’ bodies in XR environments. While previous methods have relied on color cues, depth data, or RGB-based deep learning models, these approaches often su”er from false positives and limited user distance perception. Methods In this work, we introduce a video-based self-avatar system pipeline that leverages RGB-D information to overcome these limitations. We explore two state-of-the-art real time segmentation networks, AsymFormer and ThunderNet, and fine-tune them with the RGB-D Segment Egocentric Bodies dataset, a purpose-built collection of over 9,000 RGB–depth–label triplets. We describe our complete pipeline, from dataset creation and network training to its integration into a Unity- based XR engine. Results To evaluate our approach, we conducted both quantitative and usercentered studies. Quantitatively, the AsymFormer networked trained with our custom RGB-D dataset achieves up to a 14.3% improvement in segmentation accuracy compared with former egocentric real time body segmentation, while significantly reducing false positives. The qualitative evaluation involved two steps: a video-based assessment and a mixed-reality (MR) user study. Participants in the video-based evaluation perceived substantial quality improvements when using RGB-D segmentation, reporting a 39.6% gain over the RGB-only baseline. The MR evaluation further revealed that RGB-D–based self-avatars enhance users’ distance perception. Nevertheless, participants also reported noticeable latency introduced by the RGB-D processing pipeline, especially when compared with a simpler depth-only approach. Discussion Overall, our findings demonstrate the advantages of RGB-D information for video-based self-avatar generation while highlighting current challenges such as latency and hardware limitations. Dataset is available at https://huggingface.co/datasets/ExtendedRealityLab/RGB-D-SegmentEgocentricBodies .
Title: RGB-D egocentric segmentation of human bodies for XR applications
Description:
Introduction Video-based self-avatars represent a promising approach for displaying users’ bodies in XR environments.
While previous methods have relied on color cues, depth data, or RGB-based deep learning models, these approaches often su”er from false positives and limited user distance perception.
Methods In this work, we introduce a video-based self-avatar system pipeline that leverages RGB-D information to overcome these limitations.
We explore two state-of-the-art real time segmentation networks, AsymFormer and ThunderNet, and fine-tune them with the RGB-D Segment Egocentric Bodies dataset, a purpose-built collection of over 9,000 RGB–depth–label triplets.
We describe our complete pipeline, from dataset creation and network training to its integration into a Unity- based XR engine.
Results To evaluate our approach, we conducted both quantitative and usercentered studies.
Quantitatively, the AsymFormer networked trained with our custom RGB-D dataset achieves up to a 14.
3% improvement in segmentation accuracy compared with former egocentric real time body segmentation, while significantly reducing false positives.
The qualitative evaluation involved two steps: a video-based assessment and a mixed-reality (MR) user study.
Participants in the video-based evaluation perceived substantial quality improvements when using RGB-D segmentation, reporting a 39.
6% gain over the RGB-only baseline.
The MR evaluation further revealed that RGB-D–based self-avatars enhance users’ distance perception.
Nevertheless, participants also reported noticeable latency introduced by the RGB-D processing pipeline, especially when compared with a simpler depth-only approach.
Discussion Overall, our findings demonstrate the advantages of RGB-D information for video-based self-avatar generation while highlighting current challenges such as latency and hardware limitations.
Dataset is available at https://huggingface.
co/datasets/ExtendedRealityLab/RGB-D-SegmentEgocentricBodies .

Related Results

Composing egocentric and allocentric maps for flexible navigation
Composing egocentric and allocentric maps for flexible navigation
Abstract Egocentric representations of the environment have historically been relegated to being used only for simple forms of spatial behaviour such as stimulus-re...
A SAM2-Driven RGB-T Annotation Pipeline with Thermal-Guided Refinement for Semantic Segmentation in Search-and-Rescue Scenes
A SAM2-Driven RGB-T Annotation Pipeline with Thermal-Guided Refinement for Semantic Segmentation in Search-and-Rescue Scenes
High-quality RGB–thermal infrared (RGB-T) semantic segmentation datasets are crucial for search-and-rescue (SAR) applications, yet their development is hindered by the scarcity of ...
RGB versus Early-Fusion RGB-D Glass Segmentation
RGB versus Early-Fusion RGB-D Glass Segmentation
Transparent glass is a persistent perception hazard for indoor mobile robots: RGB boundaries can be visually ambiguous, while commodity depth sensors often return missing or distor...
A SAM2-Driven RGB-T Annotation Pipeline with Thermal-Guided Refinement for Semantic Segmentation in Search-and-Rescue Scenes
A SAM2-Driven RGB-T Annotation Pipeline with Thermal-Guided Refinement for Semantic Segmentation in Search-and-Rescue Scenes
High-quality RGB–thermal infrared (RGB-T) semantic segmentation datasets are crucial for search-and-rescue (SAR) applications, yet their development is hindered by the scarcity of ...
Multiple surface segmentation using novel deep learning and graph based methods
Multiple surface segmentation using novel deep learning and graph based methods
<p>The task of automatically segmenting 3-D surfaces representing object boundaries is important in quantitative analysis of volumetric images, which plays a vital role in nu...
AI‐enabled precise brain tumor segmentation by integrating Refinenet and contour‐constrained features in MRI images
AI‐enabled precise brain tumor segmentation by integrating Refinenet and contour‐constrained features in MRI images
AbstractBackgroundMedical image segmentation is a fundamental task in medical image analysis and has been widely applied in multiple medical fields. The latest transformer‐based de...
Depth-aware salient object segmentation
Depth-aware salient object segmentation
Object segmentation is an important task which is widely employed in many computer vision applications such as object detection, tracking, recognition, and ret...

Back to Top