Javascript must be enabled to continue!
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
View through CrossRef
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose visionlanguage models. When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model's initial predictions and incrementally zooms into subregions to refine the grounding result. This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations. We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning.
Institute of Electrical and Electronics Engineers (IEEE)
Title: DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
Description:
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language.
In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization.
Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose visionlanguage models.
When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model's initial predictions and incrementally zooms into subregions to refine the grounding result.
This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations.
We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning.
Related Results
Performance Analysis and Optimization Designs for HVDC Grounding Electrodes
Performance Analysis and Optimization Designs for HVDC Grounding Electrodes
High voltage direct current(HVDC) grounding electrodes can provide a current-leakage channel for ground faults, ensuring HVDC systems’ safety and reliability. The HVDC grounding el...
Grounding of the mobile radio nodes in soils with vertical differential electrophysical causes
Grounding of the mobile radio nodes in soils with vertical differential electrophysical causes
When using overestimated values of soil parameters, the calculated resistance of the grounding electrodes will be increased and a larger number of electrodes will be required. This...
Provocative Tests in Diagnosis of Thoracic Outlet Syndrome: A Narrative Review
Provocative Tests in Diagnosis of Thoracic Outlet Syndrome: A Narrative Review
Abstract
Thoracic outlet syndrome (TOS) is a group of conditions caused by the compression of the neurovascular bundle within the thoracic outlet. It is classified into three main ...
Look Before You Leap: Context-Sensitive GUI Grounding for Boosting Automated Extended Reality (XR) Testing
Look Before You Leap: Context-Sensitive GUI Grounding for Boosting Automated Extended Reality (XR) Testing
In recent years, Extended Reality (XR) has emerged as a transformative technology, offering users immersive and interactive experiences across diversified virtual or virtual-real e...
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Task
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Task
Domain adaptation, particularly through grounding techniques, is widely adopted for multimodal large language models (MLLMs) to enhance the perception of graphical user interface (...
PENGARUH KELEMBABAN TANAH TERHADAP TAHANAN PENTANAHAN STUDI KASUS PADA GARDU INDUK KEMAYORAN 150 kV
PENGARUH KELEMBABAN TANAH TERHADAP TAHANAN PENTANAHAN STUDI KASUS PADA GARDU INDUK KEMAYORAN 150 kV
Abstract
The value of grounding resistance at the substation should be 0 Ω or less than 1 Ω. The value of grounding resistance is influenced by the resistivity and the ground...
Grounding Performance of Hydrogel, Silica Gel and Charcoal Ash as Additive Material in Grounding System
Grounding Performance of Hydrogel, Silica Gel and Charcoal Ash as Additive Material in Grounding System
Grounding enhancement materials (GEMs) are one of the additive materials which can change the grounding performance without lots of significant costs. The study aimed to assess the...
Characteristics and processes of registered nurses’ clinical reasoning and factors relating to the use of clinical reasoning in practice: a scoping review
Characteristics and processes of registered nurses’ clinical reasoning and factors relating to the use of clinical reasoning in practice: a scoping review
Objective:
The objective of this review was to examine the characteristics and processes of clinical reasoning used by registered nurses in clinical practice, and to id...

