Javascript must be enabled to continue!
RGB versus Early-Fusion RGB-D Glass Segmentation
View through CrossRef
Transparent glass is a persistent perception hazard for indoor mobile robots: RGB boundaries can be visually ambiguous, while commodity depth sensors often return missing or distorted measurements. This study isolates the effect of adding raw depth to a lightweight transformer segmenter. Two SegFormer B0 models were trained on identical partitions of the 3,009-image RGB-D Glass Surface Detection dataset using the same optimization and checkpoint-selection protocol: a three-channel RGB model and a four channel early-fusion RGB-D model. On the fixed 609-image test split, RGB outperformed RGB-D on seven of eight reported measures, reaching 0.9448 pixel accuracy, 0.8234 mean intersection over union, 0.7108 glass IoU, 0.8376 recall, and 0.8310 F1. RGB-D retained higher precision (0.8374 versus 0.8244). RGB also reduced mean inference latency by 10.7% (10.24 versus 11.47 ms) and increased throughput by 12.0% (97.63 versus 87.18 frames/s). The frozen checkpoints were then evaluated in 20 Pioneer P3-DX navigation runs. At the selected center-ratio threshold of 0.35, each model produced glass-triggered stops in both of its two glass runs (4/4 combined). Across all recorded thresholds, RGB triggered in 5/5 glass runs and RGB-D in 3/5; the independent distance safety stop terminated the remaining two RGB-D runs. Neither model produced a glass trigger in the ten non-glass runs. Under this controlled early-fusion design, raw depth did not improve unseen-test segmentation and reduced threshold robustness during robot deployment.
Title: RGB versus Early-Fusion RGB-D Glass Segmentation
Description:
Transparent glass is a persistent perception hazard for indoor mobile robots: RGB boundaries can be visually ambiguous, while commodity depth sensors often return missing or distorted measurements.
This study isolates the effect of adding raw depth to a lightweight transformer segmenter.
Two SegFormer B0 models were trained on identical partitions of the 3,009-image RGB-D Glass Surface Detection dataset using the same optimization and checkpoint-selection protocol: a three-channel RGB model and a four channel early-fusion RGB-D model.
On the fixed 609-image test split, RGB outperformed RGB-D on seven of eight reported measures, reaching 0.
9448 pixel accuracy, 0.
8234 mean intersection over union, 0.
7108 glass IoU, 0.
8376 recall, and 0.
8310 F1.
RGB-D retained higher precision (0.
8374 versus 0.
8244).
RGB also reduced mean inference latency by 10.
7% (10.
24 versus 11.
47 ms) and increased throughput by 12.
0% (97.
63 versus 87.
18 frames/s).
The frozen checkpoints were then evaluated in 20 Pioneer P3-DX navigation runs.
At the selected center-ratio threshold of 0.
35, each model produced glass-triggered stops in both of its two glass runs (4/4 combined).
Across all recorded thresholds, RGB triggered in 5/5 glass runs and RGB-D in 3/5; the independent distance safety stop terminated the remaining two RGB-D runs.
Neither model produced a glass trigger in the ten non-glass runs.
Under this controlled early-fusion design, raw depth did not improve unseen-test segmentation and reduced threshold robustness during robot deployment.
Related Results
Effects of thrombospondin-1 (THBS1) and toll-like receptor 4 (TLR-4) levels in the serum and synovial fluid of patients with knee osteoarthritis
Effects of thrombospondin-1 (THBS1) and toll-like receptor 4 (TLR-4) levels in the serum and synovial fluid of patients with knee osteoarthritis
Background: To explore the changes in and significance of Thrombospondin-1 (THBS1) and Toll-like receptor 4 (TLR-4) levels in the serum and joint fluid of patients with knee osteoa...
Unveiling the third dimension of glass
Unveiling the third dimension of glass
Glass as a material has always fascinated architects. Its inherent transparency has given us the ability to create diaphanous barriers between the interior and the exterior that al...
The Nuclear Fusion Award
The Nuclear Fusion Award
The Nuclear Fusion Award ceremony for 2009 and 2010 award winners was held during the 23rd IAEA Fusion Energy Conference in Daejeon. This time, both 2009 and 2010 award winners w...
Determinants of lateral fusion in single-level oblique lateral lumbar interbody fusion: a retrospective analysis of fusion patterns and clinical outcomes
Determinants of lateral fusion in single-level oblique lateral lumbar interbody fusion: a retrospective analysis of fusion patterns and clinical outcomes
Study Design: Retrospective cohort study.Purpose: This study aimed to (1) determine the incidence of lateral fusion following single-level oblique lateral interbody fusion (OLIF); ...
RGB-D egocentric segmentation of human bodies for XR applications
RGB-D egocentric segmentation of human bodies for XR applications
Introduction
Video-based self-avatars represent a promising approach for displaying users’ bodies in XR environments. While previous methods have relied on colo...
A SAM2-Driven RGB-T Annotation Pipeline with Thermal-Guided Refinement for Semantic Segmentation in Search-and-Rescue Scenes
A SAM2-Driven RGB-T Annotation Pipeline with Thermal-Guided Refinement for Semantic Segmentation in Search-and-Rescue Scenes
High-quality RGB–thermal infrared (RGB-T) semantic segmentation datasets are crucial for search-and-rescue (SAR) applications, yet their development is hindered by the scarcity of ...
Fusión de Datos: Imputación y Validación
Fusión de Datos: Imputación y Validación
Las actitudes, el conocimiento y las acciones generalmente se basan en muestras. Algunos basan sus conclusiones en muestras pequeñas y pocas veces toman en cuenta la magnitud de lo...
AI‐enabled precise brain tumor segmentation by integrating Refinenet and contour‐constrained features in MRI images
AI‐enabled precise brain tumor segmentation by integrating Refinenet and contour‐constrained features in MRI images
AbstractBackgroundMedical image segmentation is a fundamental task in medical image analysis and has been widely applied in multiple medical fields. The latest transformer‐based de...

