Welcome to the 14th issue of Agents in Practice.
This week I am in San Francisco attending COLM 2026, so I will introduce a couple COLM papers that seemed interesting. Please email me if you are in this area and want to meet up! I will be here until Friday night.
📬 Agents in Practice is a weekly newsletter on agentic AI research and applications. Subscribe here →
Structured reasoning in chain-of-thought

Researchers from Seoul National University and Carnegie Mellon University introduce Cognitive Chain-of-Thought (CoCoT) which defines structured stages to the reasoning of VLMs through prompts. It consists of extracting grounded facts (Perception), inferring situations (Situation), then applying social norms (Norm). Experiments show that giving such structure to reasoning improves performance on various datasets including VAGUE (Intent Disambiguation) and MoMentS (Theory-of-Mind), even when naive reasoning performed worse than no reasoning. The authors also show that traces from CoCoT can be used to fine-tune models to “internalize” the structured reasoning.
Personal Thoughts
Recent models show amazing performance, but users also complain that the reasoning takes forever and goes “out of hand”. Such structured reasoning may allow for more concentrated effort by the model, and also limit the agent’s reasoning to a usable boundary. It would be interesting to investigate if CoCoT results in less or more reasoning. I would also be curious to see how CoCoT performs on non-cognitive tasks, and whether a more general structured reasoning could be developed.
Read more
How tools shape agent behavior

Researchers from Purdue, Microsoft, and UChicago compare tool architectures that allow for similar information and actions on different interfaces and observe the impact on agent behavior. They compare 6 different tool architectures:
- Baseline BashOnly, allowing only the Bash interface
- Atomic, providing common shell actions as tools
- NLSearch, that adds a natural-language search tool for code snippets
- Python, which replaces tools with Python code execution
- HypoTrack, which allows the agent to record hypotheses
- Scratchpad, which allows the agent to record free-form thinking
They find that Atomic is the only method that improves pass^N over the baseline for all three models in the main experiment. The authors suggest this is likely due to reduced interaction errors due to tools being simple. However, when complex interactions are needed, Atomic requires more tool calls, making the run more expensive, while Python allows more work in a single interaction. On the exploration side, the authors find that NLSearch does improve exploration, but has no noticeable impact on solution diversity. The scaffolding methods HypoTrack and Scratchpad are shown to have limited effect.
Personal Thoughts
Although less “AI-flashy” than model training, building the tools and harness around the agents alters the behavior and performance of the agent meaningfully. Powerful tools like Bash and Python allow the agent to do anything, which enables the agent to make complex calls and try multiple different methods if one approach fails. However, such tools can be a significant security risk, and often cause errors in tool calls, resulting in additional cost. For the best of both worlds, it seems wise to start by allowing Bash / Python access to agents in a development environment, then extract commonly used commands or code as tools, and only allow those in production deployment.
Read more
Personal Anecdote: Subjectivity in evaluating traces

I recently got a paper accepted to the AIWolfDial workshop at INLG 2026. The paper tests different ways to make agents more consistent between what they say and what they do in a multi-agent setting. I defined sentence-action inconsistency as an agent promising to do something and not doing it, or promising not to do something and doing it anyway, and counted the examples from each game of AIWolfDial. However, as I reviewed the agent trajectories, I found myself often having to make subjective decisions. An agent may say “I will not act unless more evidence is present,” then act after another agent says something new. Does that utterance count as new evidence, or did the agent break its promise? Different reviewers could reasonably disagree. For this paper, I was the sole reviewer that annotated every inconsistency after an initial LLM pass, but for larger-scale evaluation I could see different reviewers making different decisions. Before scaling up the annotation, I would have a couple of reviewers independently annotate the same traces, compare their disagreements, and refine the definition.
One-liners
- Mistral released a public preview of Mistral Large 4 (Oct 6) with weights promised by the end of October.
- Anthropic released Claude Haiku 5.5 (Oct 7), adding adjustable reasoning effort and targeting cheaper compaction, summarization, and coding subagents.
- Following Jev, which I mentioned last week, more options for structured decision-making are appearing: Cloudflare’s Clef and Clef-flash (Oct 1), AWS’s Strands Decider 2B (Oct 1), OpenAI’s Decisions API (Oct 6), and Liquid AI’s d1-3B and d1-omni-600M (Oct 7).
- llama.cpp added support for decision models, exposing a Jev-compatible API for running structured decisions locally.
Subscribe to Agents in Practice
Get new issues of Agents in Practice — a newsletter summarizing exciting new research and applications in agentic AI — delivered to your inbox.