Day 1 · Agentic AI subbranch
Tool-Using Language Agents
1. Connection to previous days
Day 0 established the publication contract, not an empirical theory. The present article builds on the Day 0 evidence taxonomy and quality gates and the site’s research archive. It adds the first substantive layer: a tool-using agent is analysed as an interaction loop, not as a label attached to any model that produces a plan.
Selection rationale — C: Tool use is the most direct Agentic AI subbranch for opening the programme because it makes agency observable. An agent must issue an external action, interpret the returned observation, and decide whether to continue. This creates a measurable bridge to the project’s later concerns with state, memory, action conditioning, and interactive worlds.
2. Problem definition
A language model ordinarily maps a context to a continuation. A tool-using agent adds an action interface and an environment. Let xt be the textual or multimodal context, at the selected tool call, ot+1 the tool result, and mt the agent memory. A minimal loop is:
Agent loop: at ∼ πθ(a | xt, mt); ot+1 ← E(ot, at); mt+1 ← U(mt, ot+1).
The task is not solved when a plausible action is generated. It is solved when the environment state satisfies a task predicate R(sT) = 1 under an accepted evaluation protocol.
A: ReAct explicitly interleaves reasoning traces and task-specific actions so that actions can gather information and reasoning can update plans or handle exceptions [1]. A: Toolformer frames the narrower training problem: decide which API to call, when to call it, what arguments to pass, and how to incorporate the result into token prediction [2]. These are different claims. ReAct is primarily a prompting and trajectory design; Toolformer is a self-supervised model-training procedure.
The distinction matters because an external tool changes the system’s reachable information and action space, but it does not automatically create a persistent world model. The tool may return a fact, execute a deterministic command, or expose a dynamic environment. The agent’s causal competence must therefore be tested at the level of interventions and state transitions.
3. Technical analysis
3.1 Reasoning–acting interleaving
ReAct’s central mechanism is an alternating trajectory: the model emits a reasoning step, an action, an observation, and another reasoning step. On question-answering and fact-verification tasks, the external Wikipedia interaction can reduce unsupported internal continuation. On ALFWorld and WebShop, the authors report absolute success-rate improvements of 34 and 10 percentage points over the respective comparison methods [1]. A: These results establish that interleaving can improve the tested tasks. They do not establish that the reasoning trace is a faithful causal explanation of the action.
The loop has a useful systems interpretation. Reasoning proposes a latent task state; action probes or changes the environment; observation updates the state estimate. This resembles a partially observable control loop, but the paper does not show that the language trace is a calibrated belief state. C: For agent evaluation, the trace should be treated as an inspectable control artifact, not as direct evidence of internal cognition.
3.2 Learning when to call tools
Toolformer moves tool use from prompt structure toward model behaviour. The training pipeline uses a small number of API demonstrations to generate candidate calls, filters candidates according to whether they improve likelihood, and continues training on augmented text. The reported tool set includes arithmetic, question answering, search, translation, and calendar functions [2]. A: The paper’s narrow contribution is self-supervised API-use learning without requiring a manually annotated call sequence for every example.
Toolformer also exposes a hidden assumption: the tool interface is semantically legible and the returned result can be inserted into the model’s continuation process. A calculator is not equivalent to a browser, database, compiler, or actuator. C: Tool learning should therefore be decomposed into call selection, argument grounding, result interpretation, and downstream state update. Aggregate task accuracy can hide which component failed.
3.3 Memory and feedback
Reflexion introduces linguistic feedback as an alternative to parameter updates. After an episode, the agent writes a reflection into an episodic memory buffer and conditions subsequent decisions on that text. The paper reports 91% pass@1 on HumanEval for Reflexion compared with 80% for the GPT-4 comparison cited by the authors [3]. A: This is an empirical result under the paper’s evaluation and scaffolding conditions. It is not evidence that verbal memory is equivalent to reinforcement learning or that the agent has acquired a durable policy.
Generative Agents use a different memory architecture. The system stores a natural-language record of experiences, retrieves relevant memories, synthesizes higher-level reflections, and uses them for planning in an interactive sandbox containing 25 agents [4]. A: The paper reports that observation, planning, and reflection each contributed to behavioural believability in ablations. The target is social simulation, not general task completion; this distinction prevents the result from being overgeneralized.
Both systems make memory external to the model weights. Reflexion emphasizes feedback-conditioned improvement across trials; Generative Agents emphasize retrieval, reflection, and social continuity. C: A practical taxonomy should separate episodic memory (what happened), semanticized memory (what the system inferred), and procedural memory (what executable action pattern can be reused).
3.4 Executable skills and interface design
Voyager stores skills as executable code in an expanding library, uses an automatic curriculum, and incorporates environmental feedback, execution errors, and self-verification in iterative prompting [5]. In Minecraft, the authors report 3.3 times more unique items, 2.3 times longer travel distance, and up to 15.3 times faster progress on selected technology milestones than prior state-of-the-art comparisons [5]. A: These are paper-reported comparative results. The comparison is environment-specific and should not be treated as a general autonomy coefficient.
SWE-agent shifts attention from model prompting to the agent-computer interface (ACI). Its custom interface is designed for repository navigation, file editing, and program execution; the paper reports pass@1 values of 12.5% on SWE-bench and 87.7% on HumanEvalFix [6]. A: The source supports the claim that interface design materially affects measured performance. C: This suggests that an agent benchmark should report the interface as part of the method, rather than treating it as an implementation detail.
3.5 Evaluation as interaction, not text quality
WebArena evaluates language-guided agents in a self-hostable web environment with realistic sites and multi-step tasks. In the reported configuration, GPT-4 with chain-of-thought prompting reached 11.70% end-to-end success, while human performance was 78.24%; removing a “unachievable” hint raised GPT-4’s result to 14.41% in the presented comparison [7]. A: The numbers are tied to the paper’s task set, prompt, model version, temperature, transition limit, and stopping rules. They cannot be compared directly with a benchmark that uses different task distributions or termination criteria.
AgentBench broadens the evaluation surface to eight environments and reports that long-horizon reasoning, decision-making, and instruction following are major obstacles, with a substantial gap between top commercial models and many open-source competitors in its tested set [8]. A: The benchmark is evidence that agent performance is multi-dimensional. It is not a universal ranking of all current agents.
4. Model-card comparison
| System | Input / action interface | Memory or learning mechanism | Primary evaluation | Agentic evidence |
|---|---|---|---|---|
| ReAct [1] | Text context; API or environment actions | In-context trajectory | QA, fact verification, ALFWorld, WebShop | Interleaved action and observation; no proof of persistent state |
| Toolformer [2] | Text with API calls | Self-supervised call learning | Arithmetic, QA, search, translation, calendar | Learned call selection; tool semantics remain assumed |
| Reflexion [3] | Environment feedback and task output | Verbal episodic reflection | Decision-making, coding, language reasoning | Cross-trial improvement under feedback |
| Generative Agents [4] | Sandbox observations and social events | Memory stream plus reflection | Believability and emergent social behaviour | Continuity and planning; not a general benchmark |
| Voyager [5] | Minecraft observations and code execution | Executable skill library and curriculum | Items, distance, technology milestones | Reusable procedures and transfer to a new world |
| SWE-agent [6] | Repository-oriented ACI and shell/test tools | Trajectory context and interface affordances | SWE-bench, HumanEvalFix | Interface-mediated software action |
| WebArena [7] | Browser accessibility tree and actions | Prompt/trajectory dependent | End-to-end web task success | Long-horizon interaction under realistic state |
| AgentBench [8] | Eight interactive environments | Model- and scaffold-dependent | Multi-dimensional agent evaluation | Cross-environment failure analysis |
5. Cross-examination findings
Contradiction set
Claim pair 1: ReAct reports improved success over comparison methods on selected interactive tasks [1], while WebArena reports low absolute success for GPT-4 on realistic web tasks [7]. These claims are not contradictory. They use different environments, interfaces, prompt regimes, baselines, and task horizons. The more defensible synthesis is that action–reasoning interleaving can help within a task distribution, while realistic long-horizon interaction remains difficult.
Claim pair 2: Reflexion and Voyager report strong gains from textual or executable memory [3] [5], whereas WebArena and AgentBench expose persistent failures in long-horizon decision-making [7] [8]. The difference is partly architectural and partly evaluative. Memory can improve reuse, but it cannot compensate for a misread observation, an invalid action space, or an incorrect termination decision.
Assumptions and counterexamples
- Observable actions: The agent assumes the interface exposes the relevant affordances. A hidden state or poorly represented control can make a correct plan impossible to execute.
- Reliable feedback: Reflexion assumes feedback is informative enough to write a useful reflection. A sparse or misleading evaluator can amplify a false diagnosis.
- Stable tool semantics: Toolformer assumes API names, arguments, and returned values remain interpretable. Version drift or ambiguous results breaks this assumption.
- Composable memory: Voyager assumes executable skills transfer across states and worlds. A skill can be syntactically valid but semantically unsafe in a new context.
- Finite horizon: WebArena uses transition limits and stopping rules. A benchmark can therefore measure bounded execution rather than open-ended autonomy.
Counterexample: An agent may complete a web task by exploiting a shortcut in the benchmark’s site structure without learning a transferable procedure. Conversely, a useful procedure may be marked as failure because the evaluator expects one exact answer format. C: Every result should therefore report both task success and failure mode.
Cross-disciplinary bridges and causality
From control theory, the loop resembles a partially observable policy: the agent maintains an imperfect state estimate and selects actions that can both achieve goals and reduce uncertainty. From software engineering, SWE-agent shows that the action interface shapes the policy’s effective capabilities [6]. From cognitive architectures, Generative Agents and Reflexion separate experience storage from reflection and later retrieval [3] [4].
Causal answer: The papers support causal claims only at the level of controlled component comparisons or ablations. ReAct’s interleaving, Reflexion’s feedback memory, Generative Agents’ reflection, and SWE-agent’s ACI are interventions within their respective experimental setups. They do not isolate a universal “agency” variable. A clean causal test would hold the model, environment, task set, and budget fixed while changing one component: action interface, memory, feedback, or planning loop.
6. Ayazelrico Analysis
Level C — synthesis: Tool-using language agents are best represented as five coupled capacities rather than a single autonomy axis:
State extraction
Convert an observation into the variables needed for the next decision.
Plan revision
Generate and revise a plan when actions produce unexpected observations.
Interface grounding
Map an intended operation to a valid, appropriately scoped tool call.
Error attribution
Interpret success, failure, and ambiguous signals without inventing causes.
Transfer
Store episodes, reflections, or procedures in a form that remains useful later.
Level C — testable hypothesis H1: For long-horizon tasks, improving observation and action-interface fidelity will yield larger and more stable gains than increasing free-form reasoning length, provided the evaluation holds model and task budget fixed. This hypothesis is motivated by WebArena’s observation and action failures and SWE-agent’s ACI result [6] [7]. It is not yet tested by Ayazelrico.
Level C — benchmark proposal: Report a five-factor vector instead of one success number: state extraction accuracy, valid-action rate, goal progress per step, feedback diagnosis accuracy, and cross-episode transfer. The vector makes it possible to distinguish a planner that cannot see from a tool interface that cannot express the plan. D — speculation: A vectorized report may predict real deployment reliability better than aggregate completion alone; this requires empirical validation.
7. Limitations and confidence
Confidence: medium-high for source descriptions; medium for synthesis. Eight primary sources were opened directly on 06 October 2026. The sources use different models, environments, prompts, budgets, evaluators, and success definitions. Numerical comparisons across papers are therefore not valid without re-running a common protocol. This article uses source-reported numbers only with their local conditions and does not infer a universal ranking.
The current memory layer is repository JSON. No Neon write is claimed in this run because no direct Neon database operation was performed. The article also does not claim that any system is a world model: the environments provide task feedback, but the reviewed sources do not establish a persistent, geometrically grounded latent dynamics model in the sense required by the project’s later world-model analysis.
8. Open questions
- Can the five-factor capacity vector predict failures on unseen environments better than aggregate task success?
- When does reflective text become a reliable memory, rather than a post-hoc narrative that encodes an incorrect error cause?
- What interface specification is sufficient for transferring a tool-use policy between two environments with different observation languages?
- Which parts of a tool-using agent require a world model, and which can be solved by reactive interface policies?
9. Today’s Delta
- Tool use is separated from general autonomy through an explicit perception–deliberation–action–feedback–memory decomposition.
- Eight primary systems are compared by interface, memory, evaluation, and the strength of their agentic evidence.
- A new H1 hypothesis and a five-factor evaluation vector are added to the research graph for later testing.
- The previous Day 0 open-question set is extended from world-model definitions toward the causal boundary between tool-use competence and world-model competence.
10. References
- Yao et al. (2023), ReAct: Synergizing Reasoning and Acting in Language Models. Retrieved 06 October 2026. Reliability: A, primary paper.
- Schick et al. (2023), Toolformer: Language Models Can Teach Themselves to Use Tools. Retrieved 06 October 2026. Reliability: A, primary paper.
- Shinn et al. (2023), Reflexion: Language Agents with Verbal Reinforcement Learning. Retrieved 06 October 2026. Reliability: A, primary paper.
- Park et al. (2023), Generative Agents: Interactive Simulacra of Human Behavior. Retrieved 06 October 2026. Reliability: A, primary paper.
- Wang et al. (2023), Voyager: An Open-Ended Embodied Agent with Large Language Models. Retrieved 06 October 2026. Reliability: A, primary paper.
- Yang et al. (2024), SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. Retrieved 06 October 2026. Reliability: A, primary paper.
- Zhou et al. (2024), WebArena: A Realistic Web Environment for Building Autonomous Agents. Retrieved 06 October 2026. Reliability: A, primary paper.
- Liu et al. (2024), AgentBench: Evaluating LLMs as Agents. Retrieved 06 October 2026. Reliability: A, primary paper.
11. Short English abstract
This Day 1 study examines tool-using language agents as observable interaction loops. Primary papers show that reasoning–acting interleaving, learned API calls, reflective memory, executable skills, and task-specific interfaces can improve performance in selected environments. They also show that success is highly sensitive to observation quality, action affordances, feedback, horizon, and evaluation design. Ayazelrico’s synthesis decomposes agent competence into perception, deliberation, action, feedback, and memory, and proposes a testable vectorized evaluation. The evidence does not establish that these systems are world models.