SLICC ยท benchmark results

Browser Agent Benchmark: Claude vs. GPT

How well does SLICC do its work with different models behind it? We ran SLICC's agent with 11 model configurations from Anthropic, OpenAI and Moonshot AI on BU Bench V2.1, the browser agent benchmark published by Browser Use. Each run hands the agent one long, multi-step web task. An LLM judge then scores the result against the task's weighted rubric, and the runner records what the run cost and how long it took.

The numbers on this page come from the snapshot of 29 September 2026, 08:55 UTC: 1,315 runs covering 191 of the benchmark's 200 tasks. Not every configuration ran every task, so each score shows the number of judged runs behind it. The numbers will change as more runs come in; the updates below say what changed and when.

Recent updates
What changed in the results, newest first.
2026-09-29
Claude Sonnet 5.5 scores level with Claude Opus 5.5 at a lower cost

Claude Sonnet 5.5, released on 28 September, scores level with Claude Opus 5.5 at a lower cost. On the 172 tasks where both have a judged run, Sonnet 5.5 scored 50 and Opus 5.5 scored 51, and Opus 5.5 cost 18% more per task. At low effort, Sonnet 5.5 scored 34 at $0.170 per task. That is a higher score at a lower cost than GPT-6 Luna (27 at $0.365), so GPT-6 Luna is no longer on the score-vs-cost frontier.

Pending: Sonnet 5.5 runs for 10 tasks, at both effort levels, are not in this snapshot. They belong to one batch of the run (shard 12) that failed during setup and was started again on 29 September. Both Sonnet 5.5 rows are marked as pending until those runs are in.

Ranking and score vs cost

Score is the mean rubric score times 100: the share of rubric weight the judge marked as met, averaged over judged runs. It is not the share of tasks solved. Cost per task is the mean spend of the runs that finished, at list prices. Runs (n) is the number of judged runs behind each score.

Read the charts with the run counts in mind. Six configurations ran 178 to 191 tasks. Five ran only a 20-task screening round, every tenth task of the set, and are drawn as low-sample: Claude Fable 5.1 (19 of the 20 tasks), GPT-6 Astra, Kimi K3, Claude Sonnet 5 and GPT-6 Sol. Their scores come from a smaller and different set of tasks, and on the 19 tasks that every configuration ran, the order changes. The frontier line joins the configurations that no other configuration beats on both score and cost.

There are no confidence intervals. Most configurations ran each task once. Where Claude Opus 5.5 ran the same task three to five times, its best and worst scores on that task were about 31 points apart on average.

Model
Vendor
Effort
Score
Runs
Cost per task
Time per task
Note
Claude Opus 5.5
Anthropic
max effort
53.7
175
$5.226
2307 s
14 errored runs excluded; 49.7 if they count as 0
Claude Opus 5.5
Anthropic
default
51.0
245
$1.502
1242 s
191 tasks, plus repeat runs on the 20 screening tasks
Claude Sonnet 5.5
Anthropic
default
50.3
173
$1.218
1268 s
Pending: 10 shard-12 tasks are not in this snapshot; 178 of 191 tasks so far
Claude Fable 5.1
Anthropic
default
44.3
19
$6.380
1730 s
Screening round only (19 tasks)
GPT-6 Astra
OpenAI
default
41.5
20
$10.387
513 s
Screening round only (20 tasks)
Kimi K3
Moonshot AI
default
39.9
20
$2.837
1293 s
Screening round only (20 tasks)
Claude Sonnet 5
Anthropic
default
37.4
14
$5.575
2299 s
Screening round only (20 tasks); 6 errored runs excluded
Claude Opus 5.5
Anthropic
low effort
37.0
190
$0.428
506 s
191 tasks
Claude Sonnet 5.5
Anthropic
low effort
34.2
176
$0.170
303 s
Pending: 10 shard-12 tasks are not in this snapshot; 178 of 191 tasks so far
GPT-6 Sol
OpenAI
default
32.0
32
$3.937
910 s
Screening round only (20 tasks, up to 3 runs each)
GPT-6 Luna
OpenAI
default
27.0
207
$0.365
377 s
188 tasks, plus repeat runs on the 20 screening tasks

Head-to-head on the same tasks

These comparisons pair the runs of two configurations on the same task and count a pair only when both runs were judged; n is the number of pairs. Time and cost use the pairs where both runs finished, which can be a slightly different set. Each comparison covers its own set of tasks, so a model's score here can differ from its score in the table above. Scores use the same 0 to 100 scale. Every comparison with n below 50 involves a configuration from the screening round.

Against the older version

Against the model at the other provider

One tier up at the same provider

Thinking effort, against the same model at its default

How it was run

The task set

BU Bench V2.1 is a set of 200 long, multi-step web tasks, each scored against a weighted rubric. SLICC ran release v2.1.1. The runner decrypts the tasks in memory only, so no task text is stored in SLICC's repository.

The harness

Each task ran on a fresh SLICC instance: SLICC in Chrome on a cloud VM, with a wiped browser profile. The runner drives it from outside with the slicc command-line tool, so the agent does the task the way it would for a user. For every task the runner opens a new session with the agent's memories erased, sets the model and thinking effort, and sends the task text with a fixed instruction: do not ask clarifying questions, and end with a final answer line. It takes screenshots while the agent works and waits for the agent to finish, then records the cost and exports the transcript.

All models ran on AWS Bedrock, and every run used SLICC's built-in skills. Cost per task is SLICC's recorded spend for the agent and all its sub-agents, at list prices from pi-mono's published model list, or from models.dev where that list has no price. Time per task is how long the prompt took. From the full runs on, the runner also waits until every sub-agent is idle, plus a 2-minute quiet period, and that wait is part of the time. The screening round ran without it, so times are not fully comparable between rows.

Default effort is the model's default, adaptive thinking. Claude Opus 5.5 on Bedrock cannot turn thinking off, and low effort is the lowest level that works. Max effort is thinking level xhigh combined with effort max.

Limits per run

The judge

The judge is GPT-5.6 Luna on Bedrock. When a Luna verdict is still invalid after repair turns, GPT-5.6 Sol judges the run instead. This page does not break down how many runs each judge scored. The judge follows Browser Use's findings method: it marks each rubric item as met, violated or not assessable, and code computes the score as the met weight over the total weight. A run scores 1 when every item is met and 0 when none is. The judge prompt is read from Browser Use's repository at the pinned release, and the agent's own reasoning is never sent to the judge. Both judge models are from OpenAI, the vendor of the GPT-6 models under test.

Caveats

Attribution and data

The tasks, rubrics and judging method are from BU Bench V2.1 by Browser Use: github.com/browser-use/benchmark, release v2.1.1, commit af6c7f7, dataset SHA-256 0c014b056192e261b06f04ed6da6dcf37f2135150cf58c87ea839bd4e150e5c3. Browser Use did not run, review or endorse these results. At Browser Use's request, this page reproduces no task text, rubric or judge prompt.

The scores are SLICC's own measurements, taken from SLICC's benchmark dashboard of 29 September 2026. The BU Bench V2.1 run records behind them are not public at the time of writing. The public Hugging Face dataset ai-ecoverse/slicc-bench holds earlier runs on BU Bench V1 only and is not the source of these numbers. The benchmark runner is open source, in the packages/bench folder of the SLICC repository.

Run SLICC on your own tasks
Use the table above to pick a model and effort level for your budget.

Quick start

npx sliccy

Downloads SLICC, launches Chrome, and opens the workspace. Requires Node 22+.

Global install

npm install -g sliccy

Installs SLICC globally. Run slicc from any directory.

Desktop & Extension

macOS app: download, drag to Applications, double-click. Chrome extension: side panel agent with access to your logged-in sessions.

Open source, Apache-2.0 licensed. View on GitHub