Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Open-Rosalind: Tool-First Biomedical LLM Agents with Process-Aware Benchmarking

View through CrossRef
Abstract Large language models are increasingly used as scientific agents, yet the flexibility that benefits general-purpose agents can conflict with the accountability required in biomedical research. We study whether biomedical agents can be organized around auditable constraints rather than unconstrained autonomy. We present Open-Rosalind , a tool-first bio-agent system designed around four operational principles: evidence-grounded outputs, trace completeness, workflow-constrained execution, and explicit tool mediation for factual claims. To evaluate these principles, we introduce Open-Rosalind BioBench , a process-aware benchmark that measures not only task accuracy but also tool correctness, citation presence, trace completeness, and failure rate. On a strict in-house benchmark, the reference pipeline achieves 81.4% accuracy with complete execution traces. In multi-model ablations and paired replications, removing tools reduces accuracy by 19.3 to 26.4 percentage points, indicating that tool-first execution is the strongest and most stable contributor to performance. Constrained workflows also reduce lower-tail failures for models that are weak at free-form tool use. However, an author-independent 30-task hold-out initially revealed severe external-validity collapse on the deployment model. After diagnosing five routing and normalization failures and applying targeted fixes, hold-out accuracy improved from 17.8% to 53.3%, and the most concerning negative comparison against a no_tool baseline disappeared. These results position Open-Rosalind as a biomedical-agent study with an explicit external-validity audit, rather than as a claim that protocol constraints alone guarantee superior performance.
openRxiv
Title: Open-Rosalind: Tool-First Biomedical LLM Agents with Process-Aware Benchmarking
Description:
Abstract Large language models are increasingly used as scientific agents, yet the flexibility that benefits general-purpose agents can conflict with the accountability required in biomedical research.
We study whether biomedical agents can be organized around auditable constraints rather than unconstrained autonomy.
We present Open-Rosalind , a tool-first bio-agent system designed around four operational principles: evidence-grounded outputs, trace completeness, workflow-constrained execution, and explicit tool mediation for factual claims.
To evaluate these principles, we introduce Open-Rosalind BioBench , a process-aware benchmark that measures not only task accuracy but also tool correctness, citation presence, trace completeness, and failure rate.
On a strict in-house benchmark, the reference pipeline achieves 81.
4% accuracy with complete execution traces.
In multi-model ablations and paired replications, removing tools reduces accuracy by 19.
3 to 26.
4 percentage points, indicating that tool-first execution is the strongest and most stable contributor to performance.
Constrained workflows also reduce lower-tail failures for models that are weak at free-form tool use.
However, an author-independent 30-task hold-out initially revealed severe external-validity collapse on the deployment model.
After diagnosing five routing and normalization failures and applying targeted fixes, hold-out accuracy improved from 17.
8% to 53.
3%, and the most concerning negative comparison against a no_tool baseline disappeared.
These results position Open-Rosalind as a biomedical-agent study with an explicit external-validity audit, rather than as a claim that protocol constraints alone guarantee superior performance.

Related Results

Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract Introduction The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Evolving benchmarking practices: a review for research perspectives
Evolving benchmarking practices: a review for research perspectives
PurposeThe purpose of this study is to review a major section of the literature on benchmarking practices in order to achieve better perspectives for emerging benchmarking research...
Perceptions about benchmarking best practices among French managers: an exploratory survey
Perceptions about benchmarking best practices among French managers: an exploratory survey
PurposeThe purpose of this study is to present a discussion on the most commonly accepted benchmarking norms in the USA, the lessons learned from benchmarking experiences and see h...
Optimising tool wear and workpiece condition monitoring via cyber-physical systems for smart manufacturing
Optimising tool wear and workpiece condition monitoring via cyber-physical systems for smart manufacturing
Smart manufacturing has been developed since the introduction of Industry 4.0. It consists of resource sharing and networking, predictive engineering, and material and data analyti...
Improving SME logistics performance through benchmarking
Improving SME logistics performance through benchmarking
Purpose The purpose of this paper is to discuss the applicability of current benchmarking proposals for small and medium-sized enterprises (SMEs) and to suggest a condensed process...
Human-AI Collaboration in Clinical Reasoning: A UK Replication and Interaction Analysis
Human-AI Collaboration in Clinical Reasoning: A UK Replication and Interaction Analysis
Abstract Objective A paper from Goh et al found that a large language model (LLM) working alone outperformed American clinicians assisted...
Barriers to internal benchmarking initiatives: an empirical investigation
Barriers to internal benchmarking initiatives: an empirical investigation
PurposeThe purpose of this paper is to focus on the identification of barriers to the implementation of benchmarking initiatives. Managers have little guidance on strategies for su...
A review on benchmarking of supply chain performance measures
A review on benchmarking of supply chain performance measures
PurposeThe purpose of this paper is to redress the imbalances in the past literature of supply chain benchmarking and enhance data envelopment analysis (DEA) modeling approach in s...

Back to Top