Javascript must be enabled to continue!
Agentic Frameworks for Reasoning Tasks: An Empirical Study
View through CrossRef
<p>Recent advances in agentic frameworks have enabled the development of AI agents capable of complex reasoning and decision-making. However, systematic evidence comparing their reasoning performance, efficiency, and real-world suitability remains limited, making it difficult to select appropriate frameworks for practical use. To address this gap, we empirically evaluate 22 widely used agentic frameworks across three reasoning benchmarks: BBH, GSM8K, and ARC. We selected these frameworks from 1,200 GitHub repositories collected between January 2023 and July 2025, and developed a taxonomy based on their architectural designs. We then evaluated these frameworks under a unified experimental setting, measuring reasoning accuracy, execution time, computational cost, and the consistency of framework performance across the benchmarks. Our results show that 19 of the 22 frameworks completed all three benchmarks. </p>
<p>Among these, 12 frameworks showed consistent performance, with a mean accuracy of 74.6–75.9%, execution time of 4–6 seconds per task, and cost of 0.14–0.18¢ per task. The remaining frameworks performed worse, primarily because of orchestration issues rather than limitations in reasoning. For instance, Camel failed to complete BBH after 11 days of runtime because of uncontrolled context growth; Upsonic consumed $1,434 in a single day because extraction failures led to multiple retries, while uncontrolled context growth sharply increased prompttoken usage; and AutoGen and Mastra exhausted API quotas through iterative agent interactions that increased prompt length without improving the answers. These are system-level failures in memory management, retry policy, and context handling rather than failures in reasoning. We further found that all frameworks, including the high performers, degrade sharply regarding mathematical reasoning: among the frameworks that completed all benchmarks, mean accuracy on GSM8K was 44.35%, compared with 89.80% on BBH and 89.56% on ARC. This gap is consistent across architectures, reinforcing that current agentic designs inherit rather than overcome the base model’s weaknesses in multi-step numerical computation.</p>
<p> Overall, this study provides a systematic comparison and architectural perspective on agentic frameworks to support framework selection for reasoning-intensive software engineering tasks. Our results suggest that such selection should be guided primarily by orchestration quality—particularly memory discipline, failure handling, and cost control—rather than by architectural category or claimed reasoning capabilities alone. We have released the dataset, source code, and per-task logs to support replication and further research.</p>
<p></p>
Title: Agentic Frameworks for Reasoning Tasks: An Empirical Study
Description:
<p>Recent advances in agentic frameworks have enabled the development of AI agents capable of complex reasoning and decision-making.
However, systematic evidence comparing their reasoning performance, efficiency, and real-world suitability remains limited, making it difficult to select appropriate frameworks for practical use.
To address this gap, we empirically evaluate 22 widely used agentic frameworks across three reasoning benchmarks: BBH, GSM8K, and ARC.
We selected these frameworks from 1,200 GitHub repositories collected between January 2023 and July 2025, and developed a taxonomy based on their architectural designs.
We then evaluated these frameworks under a unified experimental setting, measuring reasoning accuracy, execution time, computational cost, and the consistency of framework performance across the benchmarks.
Our results show that 19 of the 22 frameworks completed all three benchmarks.
</p>
<p>Among these, 12 frameworks showed consistent performance, with a mean accuracy of 74.
6–75.
9%, execution time of 4–6 seconds per task, and cost of 0.
14–0.
18¢ per task.
The remaining frameworks performed worse, primarily because of orchestration issues rather than limitations in reasoning.
For instance, Camel failed to complete BBH after 11 days of runtime because of uncontrolled context growth; Upsonic consumed $1,434 in a single day because extraction failures led to multiple retries, while uncontrolled context growth sharply increased prompttoken usage; and AutoGen and Mastra exhausted API quotas through iterative agent interactions that increased prompt length without improving the answers.
These are system-level failures in memory management, retry policy, and context handling rather than failures in reasoning.
We further found that all frameworks, including the high performers, degrade sharply regarding mathematical reasoning: among the frameworks that completed all benchmarks, mean accuracy on GSM8K was 44.
35%, compared with 89.
80% on BBH and 89.
56% on ARC.
This gap is consistent across architectures, reinforcing that current agentic designs inherit rather than overcome the base model’s weaknesses in multi-step numerical computation.
</p>
<p> Overall, this study provides a systematic comparison and architectural perspective on agentic frameworks to support framework selection for reasoning-intensive software engineering tasks.
Our results suggest that such selection should be guided primarily by orchestration quality—particularly memory discipline, failure handling, and cost control—rather than by architectural category or claimed reasoning capabilities alone.
We have released the dataset, source code, and per-task logs to support replication and further research.
</p>
<p></p>.
Related Results
Non-Human Identity and Access Management for Agentic AI
Non-Human Identity and Access Management for Agentic AI
<p>Background. Identity and access management was designed for human users, and a modest population of long-lived service accounts, yet agentic artificial intelligence has in...
Agentic AI Systems: Architectures, Autonomy, and Emergent Behaviours
Agentic AI Systems: Architectures, Autonomy, and Emergent Behaviours
<p><b><i><span>Background.</span></i></b><span> Agentic artificial intelligence systems, defined by their capacity to reason, plan, ...
A Survey on Agentic AI Frameworks for Network Security
A Survey on Agentic AI Frameworks for Network Security
The development of the new category of the AI-driven systems called the Agentic AI that represents a paradigm shift in the architectural design of autonomous networked and security...
Characteristics and processes of registered nurses’ clinical reasoning and factors relating to the use of clinical reasoning in practice: a scoping review
Characteristics and processes of registered nurses’ clinical reasoning and factors relating to the use of clinical reasoning in practice: a scoping review
Objective:
The objective of this review was to examine the characteristics and processes of clinical reasoning used by registered nurses in clinical practice, and to id...
Exploring Agentic AI in Healthcare: A Study on Its Working Mechanism
Exploring Agentic AI in Healthcare: A Study on Its Working Mechanism
Introduction
Rapid advancements in artificial intelligence (AI) have ushered in an era of hyperautomation and intelligent orchestration across multiple engineer...
Logical Challenges in Artificial General Intelligence
Logical Challenges in Artificial General Intelligence
The present thesis pertains to the research area of logic for artificial intelligence (AI), and is motivated by the critical role of automated reasoning in AI, particularly by the ...
A Framework for Building Robust AI Agents
A Framework for Building Robust AI Agents
Agentic AI has evolved from experimental prototypes to a central paradigm for building intelligent systems. While early demonstrations showcase impressive autonomy through tool use...
SCALER: A Procedurally Generated, Leakage-Resistant Benchmark for Evaluating Multi-Step Reasoning in Large Language Models
SCALER: A Procedurally Generated, Leakage-Resistant Benchmark for Evaluating Multi-Step Reasoning in Large Language Models
Abstract
The evaluation of multi-step reasoning capabilities in Large Language Models (LLMs) faces three fundamental challenges that threaten the validity of curren...

