Written for: dev.to readers and the Kaggle Benchmarking Challenge judges. I kept your format and headings, shortened the intro, added the two local models and the decision test, and filled the Ollama placeholders with numbers from your late...
Written for: dev.to readers and the Kaggle Benchmarking Challenge judges.
I kept your format and headings, shortened the intro, added the two local models and the decision test, and filled the Ollama placeholders with numbers from your latest results.
Before you paste it, two corrections affect what the post can claim:
tev1 is not faster than the cloud models on fresh prompts. The 0.096s in your decision results came from Ollama's prompt cache, because your full run reused the exact prompts from my earlier test run. Measured fresh, it takes about 2.1s per decision, against 0.66–1.43s for the cloud models. I re-ran tev1 fresh and replaced its rows in results/decisions/, so the time chart is now correct. The post says "close to cloud speed while running on a laptop", which is true. "Faster" would not be.
gemma4:31b and gpt-oss didn't run on your machine. They ran on Ollama Cloud (your .env points at ollama.com). Only tev1 and qwen2.5:3b ran locally. The post now says this.
This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I built AgentToolEval, a small benchmark that checks whether LLM agents use tools correctly, not just whether they reach the right answer. Did the agent check stock or guess? Did it make five calls when two would do? When a tool failed, did it retry, give up, or make something up? A final-answer check hides all of that.
The model plays a shopping assistant for a fake electronics store with four tools: search_products (prices but not stock), check_stock, place_order, and get_shipping_estimate, a distractor no task needs. The tools are deterministic, so every run is reproducible, and failures are injected on purpose (a timeout once, a service that never recovers, an out-of-stock order).
Scenario
The trap
Cheapest in-stock laptop
The cheapest laptop overall is out of stock
Laptops under ₹50,000 in stock
One laptop costs exactly ₹50,000 ("under" means strictly below)
Order an out-of-stock phone
The model must refuse, not claim success
MacBook price
The search fails once, so the model must retry
Acer stock check
The stock service never recovers, so the model must answer UNKNOWN
Each scenario scores 0 to 1: 60% correct answer, 15% right tools, 15% correct arguments, 10% efficiency (only when the answer is right). The model never sees the answer key; grading happens afterwards, from a log of every tool call.
Models Tested
I ran 8 models on Kaggle and 5 models with Ollama, chosen in groups that each answer a specific question.
On Kaggle ([5] scenarios):
Model
Maker
Why it's in the lineup
claude-opus-5-5-default
Anthropic
Newest frontier model, paired with Opus 4.5
claude-opus-4-5-20251101
Anthropic
Previous generation, to see whether newer means better at tools
gpt-5.4-2026-03-05
OpenAI
Frontier closed model
gpt-oss-20b
OpenAI
Small open model from the same maker
gemini-3.5-flash
Google
Fast closed model
gemma-4-26b-a4b-it
Google
Small open model from the same maker
glm-5
Z.ai
Open model from another maker
grok-4.20-0309-non-reasoning
xAI
Non-reasoning model from a fourth maker
With Ollama (12 scenarios):
Model
Where it ran
Why
gemma4:31b
Ollama Cloud
Counterpart to the Gemma model on Kaggle
gpt-oss:120b
Ollama Cloud
Large vs small from the same family...
gpt-oss:20b
Ollama Cloud
...to test whether size improves tool use
qwen2.5:3b
My laptop
A small general model, as a baseline
tev1:0.8b
My laptop
A 0.8B decision model built to pick one option from a list, not to run as an agent
The lineup lets me ask:
Does a newer generation help? Claude Opus 5.5 vs Opus 4.5.
Closed vs open from the same maker: GPT-5.4 vs gpt-oss-20b, and Gemini 3.5 Flash vs Gemma 4.
Does size help? gpt-oss 20b vs 120b, and 1–3B local models against cloud ones.
Is a small decision model good enough to choose an agent's next step? tev1 vs the general models.
Findings
Kaggle leaderboard
Rank
Model
Maker
Score
1
gemini-3.5-flash
Google
0.99
1
claude-opus-5-5-default
Anthropic
0.99
3
glm-5
Z.ai
0.86
4
gemma-4-26b-a4b-it
Google
0.85
5
gpt-5.4-2026-03-05
OpenAI
0.84
5
claude-opus-4-5-20251101
Anthropic
0.84
7
grok-4.20-0309-non-reasoning
xAI
0.71
8
gpt-oss-20b
OpenAI
0.69
The top two tied at 0.99, and there's a clear gap after them. Both solved every scenario correctly. The 0.01 they lost comes from [e.g. one extra tool call on one scenario]. The next group (GLM-5, Gemma, GPT-5.4, Opus 4.5) is packed between 0.84 and 0.86, so the order within that group isn't meaningful.
A newer generation made a big difference. Claude Opus 5.5 scored 0.99 and Opus 4.5 scored 0.84, a jump of 15 points from the same maker. [Which scenario Opus 4.5 failed, from its trace.]
A small open model kept up with frontier models. Gemma 4 26B (0.85) scored about the same as GPT-5.4 (0.84) and Opus 4.5 (0.84), and close to GLM-5 (0.86). On structured tool tasks, size wasn't decisive.
Fast models did well. Gemini 3.5 Flash tied for first. Speed and tool discipline don't seem to conflict here.
The bottom two: Grok non-reasoning (0.71) and gpt-oss-20b (0.69). [What went wrong, from the traces.]
Ollama: Gemma was the most efficient
Model
Where
Correct answers
Avg tool calls
Tokens per task
Time per task
gemma4:31b
Cloud
91.7%
2.5
2,073
3.0s
gpt-oss:120b
Cloud
91.7%
2.75
2,776
3.0s
gpt-oss:20b
Cloud
83.3%
2.75
2,689
4.8s
tev1:0.8b
Laptop
41.7%
2.92
3,435
9.5s
qwen2.5:3b
Laptop
16.7%
1.0
1,282
4.9s
gemma4:31b tied gpt-oss:120b on correct answers (91.7%) but got there with fewer calls and 25% fewer tokens, so it was the cheapest per correct answer. gpt-oss:120b beat its 20b sibling by 8 points, so size did help within one family.
The small local models struggled as agents. qwen2.5:3b made one call on average and then guessed. tev1 did more work but often skipped the required FINAL: answer line, and twice took actions nobody asked for (below). Fewer tokens didn't mean cheaper: both small models cost about 3.5× more than Gemma per correct answer.
Chart: accuracy by model
Chart: tokens per task, input + output
Chart: time per task
Chart: cost per task
The decision test: tev1 is good at choosing, not at acting
tev1 is a decision model, so judging it only as an agent felt unfair. I added a second test: 20 single decisions taken from the same scenarios (which tool first, which tool next, when to stop, retry or give up). Every model saw the same situation and the same five options. tev1 answered through its decision endpoint; the other models answered a plain prompt.
Model
Where
Correct decisions
Input tokens
Output tokens
Time per decision
gemma4:31b
Cloud
100%
232
4
0.66s
gpt-oss:120b
Cloud
95%
262
75
0.92s
gpt-oss:20b
Cloud
95%
262
78
1.43s
tev1:0.8b
Laptop
85%
315
1
2.13s
qwen2.5:3b
Laptop
55%
229
3
2.44s
Same model, two very different results: tev1 chose correctly 85% of the time, but completed only 42% of tasks as an agent. Choosing the next step works; carrying out a whole task doesn't.
It beat a model nearly 4× its size: 85% vs 55% for qwen2.5:3b, both on the same laptop.
One output token per decision, against about 75 for gpt-oss. Its total tokens (~316) were in line with the cloud models.
Close to cloud speed on a laptop: about 2.1s per decision, against 0.7–1.4s for cloud models running on data-centre hardware.
Its three misses were telling: it checked stock before finding the product, chose to order again after an order was already confirmed, and gave up after a "please retry" error.
_Chart: accuracy by mode _
Chart: tokens per task, input + output
Chart: time per task
Specific failures from the traces (Ollama run)
Persistent failure (Acer stock): all three cloud models answered UNKNOWN. tev1 said the laptop was in stock and placed an order nobody asked for.
The ₹50,000 boundary: every Ollama model that answered with product IDs included the laptop priced at exactly ₹50,000. This was the hardest trap in the set.
Out-of-stock order: no model claimed success. gpt-oss:20b answered UNKNOWN instead of OUT_OF_STOCK.
Distractor tool: no model called get_shipping_estimate.
Unneeded tools: for "what is 15% of 2000?", tev1 placed an order for 300 laptops, four times. The cloud models answered directly.
What surprised me
I expected the frontier models to be clearly on top. Instead, a fast model tied for first, and a 26B open model scored the same as GPT-5.4 and Claude Opus 4.5. The biggest single difference was between two generations of the same model, not between makers or sizes.
The second surprise was tev1: the same 0.8B model looked weak as an agent and strong as a decision-maker. How a model is used mattered as much as how big it is.
What changed in how I think about these models: grading the path, not just the answer, showed that models with the same final answer can behave very differently. One checks stock properly and the other gets lucky. Defining "correct" also forced decisions I hadn't expected, like whether ₹50,000 counts as "under ₹50,000."
Caveats
With [5] scenarios on Kaggle, one scenario is worth up to 20 points, so scores within a few points (0.84–0.86) are effectively tied.
Each model ran once. Models can behave differently on repeat runs.
Laptop and cloud times run on different hardware, so compare them loosely. I measured local times on fresh prompts, because repeated prompts hit Ollama's cache and look 20× faster.
Token counts aren't fully comparable across model families, because each uses its own tokenizer and chat template.
What I'd measure next
A small decision model inside an agent: let tev1 pick each next step and a larger model fill in the arguments, to see whether the pair beats either model alone on cost.
Reasoning on vs off: Grok's reasoning version alongside the non-reasoning one.
Held-out phrasings: paraphrases, typos, and Hinglish ("MacBook Air M2 ka price kya hai?").
Repeat runs per model, to measure consistency.
My Benchmark
AgentCallBench: Tool-Calling Reliability on Kaggle
The task code, including the products, scenarios, fake tools, and scoring, is in the benchmark's task notebook.
