SLICC ยท benchmark results
Browser Agent Benchmark: Claude vs. GPT
How well does SLICC do its work with different models behind it? We ran SLICC's agent with 11 model configurations from Anthropic, OpenAI and Moonshot AI on BU Bench V2.1, the browser agent benchmark published by Browser Use. Each run hands the agent one long, multi-step web task. An LLM judge then scores the result against the task's weighted rubric, and the runner records what the run cost and how long it took.
The numbers on this page come from the snapshot of 29 September 2026, 08:55 UTC: 1,315 runs covering 191 of the benchmark's 200 tasks. Not every configuration ran every task, so each score shows the number of judged runs behind it. The numbers will change as more runs come in; the updates below say what changed and when.
Claude Sonnet 5.5, released on 28 September, scores level with Claude Opus 5.5 at a lower cost. On the 172 tasks where both have a judged run, Sonnet 5.5 scored 50 and Opus 5.5 scored 51, and Opus 5.5 cost 18% more per task. At low effort, Sonnet 5.5 scored 34 at $0.170 per task. That is a higher score at a lower cost than GPT-6 Luna (27 at $0.365), so GPT-6 Luna is no longer on the score-vs-cost frontier.
Pending: Sonnet 5.5 runs for 10 tasks, at both effort levels, are not in this snapshot. They belong to one batch of the run (shard 12) that failed during setup and was started again on 29 September. Both Sonnet 5.5 rows are marked as pending until those runs are in.
Ranking and score vs cost
Score is the mean rubric score times 100: the share of rubric weight the judge marked as met, averaged over judged runs. It is not the share of tasks solved. Cost per task is the mean spend of the runs that finished, at list prices. Runs (n) is the number of judged runs behind each score.
Read the charts with the run counts in mind. Six configurations ran 178 to 191 tasks. Five ran only a 20-task screening round, every tenth task of the set, and are drawn as low-sample: Claude Fable 5.1 (19 of the 20 tasks), GPT-6 Astra, Kimi K3, Claude Sonnet 5 and GPT-6 Sol. Their scores come from a smaller and different set of tasks, and on the 19 tasks that every configuration ran, the order changes. The frontier line joins the configurations that no other configuration beats on both score and cost.
There are no confidence intervals. Most configurations ran each task once. Where Claude Opus 5.5 ran the same task three to five times, its best and worst scores on that task were about 31 points apart on average.
Head-to-head on the same tasks
These comparisons pair the runs of two configurations on the same task and count a pair only when both runs were judged; n is the number of pairs. Time and cost use the pairs where both runs finished, which can be a slightly different set. Each comparison covers its own set of tasks, so a model's score here can differ from its score in the table above. Scores use the same 0 to 100 scale. Every comparison with n below 50 involves a configuration from the screening round.
Against the older version
- Claude Sonnet 5.5 against Claude Sonnet 5: 43 against 37 (+16.2%), 77% cheaper, 63% faster. n = 14.
Against the model at the other provider
- Claude Fable 5.1 against GPT-6 Astra: 44 against 43 (+2.3%), 38% cheaper, 240% slower. n = 19.
- Claude Opus 5.5 against GPT-6 Sol: 55 against 33 (+69.2%), 61% cheaper, 27% slower. n = 31.
One tier up at the same provider
- Claude Opus 5.5 against Claude Sonnet 5.5: 51 against 50 (+0.9%), 18% more expensive, 1% faster. n = 172.
- Claude Fable 5.1 against Claude Opus 5.5: 43 against 54 (-21.1%), 369% more expensive, 55% slower. n = 18.
- Claude Opus 5.5 against Claude Sonnet 5: 51 against 40 (+27.0%), 78% cheaper, 57% faster. n = 13.
- GPT-6 Astra against GPT-6 Sol: 39 against 29 (+35.7%), 163% more expensive, 42% faster. n = 14.
Thinking effort, against the same model at its default
- Claude Opus 5.5 at low effort: 37 against 52 (-28.7%), 72% cheaper, 60% faster. n = 188.
- Claude Opus 5.5 at max effort: 54 against 52 (+3.8%), 294% more expensive, 97% slower. n = 174.
- Claude Sonnet 5.5 at low effort: 34 against 51 (-33.1%), 86% cheaper, 76% faster. n = 171.
How it was run
The task set
BU Bench V2.1 is a set of 200 long, multi-step web tasks, each scored against a weighted rubric. SLICC ran release v2.1.1. The runner decrypts the tasks in memory only, so no task text is stored in SLICC's repository.
The harness
Each task ran on a fresh SLICC instance: SLICC in Chrome on a cloud VM, with a wiped browser profile. The runner drives it from outside with the slicc command-line tool, so the agent does the task the way it would for a user. For every task the runner opens a new session with the agent's memories erased, sets the model and thinking effort, and sends the task text with a fixed instruction: do not ask clarifying questions, and end with a final answer line. It takes screenshots while the agent works and waits for the agent to finish, then records the cost and exports the transcript.
All models ran on AWS Bedrock, and every run used SLICC's built-in skills. Cost per task is SLICC's recorded spend for the agent and all its sub-agents, at list prices from pi-mono's published model list, or from models.dev where that list has no price. Time per task is how long the prompt took. From the full runs on, the runner also waits until every sub-agent is idle, plus a 2-minute quiet period, and that wait is part of the time. The screening round ran without it, so times are not fully comparable between rows.
Default effort is the model's default, adaptive thinking. Claude Opus 5.5 on Bedrock cannot turn thinking off, and low effort is the lowest level that works. Max effort is thinking level xhigh combined with effort max.
Limits per run
- One hour per run, as the benchmark specifies.
- A cost cap of $10 per run. The screening batch with Claude Fable 5.1, GPT-6 Astra and Kimi K3 ran with a $30 cap.
- A run that reaches a limit is stopped and judged on the work it had done.
- The limits were not always enforced. Until a SLICC fix on 27 September, stopping a run did not stop the agent, so some runs in the screening round kept working and spending without being measured. Their recorded costs are probably too low. Some recorded runs go past the limits, up to 4,274 seconds and, for one GPT-6 Astra run, $44.879.
The judge
The judge is GPT-5.6 Luna on Bedrock. When a Luna verdict is still invalid after repair turns, GPT-5.6 Sol judges the run instead. This page does not break down how many runs each judge scored. The judge follows Browser Use's findings method: it marks each rubric item as met, violated or not assessable, and code computes the score as the met weight over the total weight. A run scores 1 when every item is met and 0 when none is. The judge prompt is read from Browser Use's repository at the pinned release, and the agent's own reasoning is never sent to the judge. Both judge models are from OpenAI, the vendor of the GPT-6 models under test.
Caveats
- Errored runs are left out of the scores. 31 runs ended in an error and 13 more have no verdict from the judge. None of them count toward a score; the 13 without a verdict still count toward time and cost. Most errors come from the harness, for example an agent that kept working after the stop signal, a transcript export that timed out or a dropped connection. The errors are not spread evenly. 14 of the 190 runs of Claude Opus 5.5 at max effort errored, 12 of them because the agent ran past its limits or kept working after the stop. Counting those errors as a score of 0 moves its mean from 0.537 to 0.497.
- Nine of the 200 tasks have no run in this snapshot: bu2-135, bu2-136, bu2-155, bu2-156, bu2-175, bu2-176, bu2-182, bu2-195 and bu2-196. They were part of the full-run plan, and the cause of the gap is not confirmed. Claude Sonnet 5.5 at both efforts also has no runs yet for 13 more tasks (the 10 pending ones and 3 others), GPT-6 Luna for 3 and Claude Opus 5.5 at max effort for 1.
- Results pool several SLICC versions. Releases shipped while the benchmark ran, so the runs span seven versions, from 6.195.0 to 6.210.1. The screening round ran before the full runs and before several harness fixes landed, and the default-effort Claude Opus 5.5 row pools runs from both periods.
- Until a SLICC fix on 28 September, any GPT-6 turn that attached a screenshot the way SLICC sent it failed on Bedrock. GPT-6 Sol and GPT-6 Astra ran only before the fix, and so did part of the GPT-6 Luna runs. The affected runs have not been counted, and the GPT-6 scores may be lower than they would be on the fixed build.
- The judge gets less evidence than in Browser Use's own runner. It receives at most 10 screenshots, where Browser Use sends up to 50. Deliverable files such as a CSV or JSON output are not passed to it as files; it sees their contents only where they appear in the transcript. SLICC also leaves the judge's reasoning effort unset, where Browser Use's runner uses xhigh.
- These are SLICC's own runs, with SLICC's harness, browser setup and judge configuration. Browser Use's published V2 charts use an earlier 60-task set, so their scores are not comparable with the scores on this page.
Attribution and data
The tasks, rubrics and judging method are from BU Bench V2.1 by Browser Use: github.com/browser-use/benchmark, release v2.1.1, commit af6c7f7, dataset SHA-256 0c014b056192e261b06f04ed6da6dcf37f2135150cf58c87ea839bd4e150e5c3. Browser Use did not run, review or endorse these results. At Browser Use's request, this page reproduces no task text, rubric or judge prompt.
The scores are SLICC's own measurements, taken from SLICC's benchmark dashboard of 29 September 2026. The BU Bench V2.1 run records behind them are not public at the time of writing. The public Hugging Face dataset ai-ecoverse/slicc-bench holds earlier runs on BU Bench V1 only and is not the source of these numbers. The benchmark runner is open source, in the packages/bench folder of the SLICC repository.
Quick start
npx sliccy
Downloads SLICC, launches Chrome, and opens the workspace. Requires Node 22+.
Global install
npm install -g sliccy
Installs SLICC globally. Run slicc from any directory.
Desktop & Extension
- https://www.sliccy.ai/download/slicc.dmg
- https://chrome.google.com/webstore/detail/akjjllgokmbgpbdbmafpiefnhidlmbgf
macOS app: download, drag to Applications, double-click. Chrome extension: side panel agent with access to your logged-in sessions.