Leaderboards
Independent product-level leaderboards
Only evaluators that test finished personal-agent products on tasks appear here. Each table is reproduced on the publisher's own scale with its sample size, evidence labels and caveats. Scores are never converted into a common unit.
United States · English
Assistant Benchmark
About this evaluator → 1–10 mean of scored dimensions. 12 products scored; Gemini Spark is listed with no scored dimensions.
- Publisher
- Assistant Benchmark (independent) (US)
- Scale
- 1–10 mean of scored dimensions
- Method
- Real accounts and real services; 15 scored dimensions anchored 1–10. Each dimension is labelled as a test or an observation. Overall = mean of the dimensions that were scored, so coverage differs by product.
- Retrieved
- 2026-10-09 · source updated 2026-10-08
- Original
- Open source leaderboard ↗
Caveats (3)
- Scores average different subsets of dimensions; compare coverage before comparing scores.
- Mixes controlled tests with usage observations.
- Single reviewer setup per product; no published intervals.
| Product | Scored | Carrying out an online task | Travel booking | Recommendation quality | Purchasing a product | Responding to emails | Proactive behavior | Running a routine | Third-party integrations | Permissions & privacy | Memory | Phone calls | Multiplayer / groups | Chained tasks | Proactive restraint | Content creation / games | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Meta Muse | 9.3 | 8/15 | 9T | 7O | 9T | 10O | 10O | — | 10O | 10O | 9O | — | — | — | — | — | — |
| Instinct | 8.5 | 12/15 | 8O | 10O | 9O | 9O | 9O | 8O | 10T | 8O | 5O | 9O | — | 9? | — | 8O | — |
| OpenAI dots | 8.4 | 7/15 | 9T | — | 9T | 6O | 10T | 8O | — | — | 8O | 9O | n/a | — | — | — | — |
| Ollie | 8.3 | 13/15 | 9T | 7T | 7T | 8T | 8T | 8T | 10T | 7T | 8T | — | n/a | 9T | 8T | 10T | 9T |
| Pally | 8.3 | 13/15 | 7T | 8T | 9T | 8T | 8T | 8T | 10T | 7T | 8T | — | n/a | 7T | 9T | 10T | 9T |
| Town | 8.3 | 7/15 | — | — | 8T | — | 8T | — | 7T | 10O | 8O | 8? | — | — | — | — | 9T |
| szn | 8.2 | 13/15 | 8T | 8T | 7T | 8T | 8T | 8O | — | 7T | 7T | 10T | 10T | — | 7T | 10O | 8T |
| Shuffle | 7.8 | 10/15 | 7.5T | 9T | 9T | 8T | 8T | — | — | 7T | 7T | 6T | n/a | n/a | 7T | — | 9T |
| Tomo | 7.7 | 11/15 | 7T | 9T | 8T | 8T | 8T | — | — | 7T | 7T | 6T | — | 9O | 7T | — | 9T |
| Asaply | 7.5 | 2/15 | — | n/a | 7T | 8T | n/a | — | — | — | — | — | n/a | — | — | — | n/a |
| Caddy | 7.3 | 9/15 | 8T | 8T | 7T | 7T | 7T | — | — | 7T | — | 7T | n/a | n/a | 7T | — | 8T |
| Grok Bot | 7.3 | 7/15 | 8O | — | 7O | 8O | 9O | — | — | 3O | — | — | — | — | 8O | — | 8O |
| Gemini Spark | — | 0/15 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Cell = source score on its 1–10 anchor scale. T = test, O = observed in use, ? = unlabelled. — = not tested, n/a = not applicable; neither is zero.
United States · English
micro1 PersonalAgentBench
About this evaluator → % of 10 workflows with trusted completion. Ranked by trusted completion, task completion breaks ties.
- Publisher
- micro1 (US)
- Scale
- % of 10 workflows with trusted completion
- Method
- 10 real-account workflows per assistant. Three expert testers per workflow; pass/fail outcome checks. One eligible attempt per assistant-workflow pair (earliest eligible, not best). Ranking by Trusted Completion, Task Completion as tie-breaker.
- Retrieved
- 2026-10-09
- Original
- Open source leaderboard ↗
Caveats (4)
- Provisional leaderboard; test dates and product versions are not stated by the publisher.
- n = 10 workflows per product: 95% intervals are wide and mostly overlap, so rank order is not statistically separated.
- Sample skews towards Google Workspace workflows; accounts and histories are not identical across products.
- 5 of 40 cells used a later attempt because the earlier attempt was ineligible.
| Source rank | Product | Trusted completion | 95% interval | Task completion | 95% interval | Workflows |
|---|---|---|---|---|---|---|
| 1 | Gemini Spark | 60% (6/10) | 31–83% | 70% (7/10) | 40–89% | 10 |
| 2 | Instinct | 40% (4/10) | 17–69% | 40% (4/10) | 17–69% | 10 |
| 3 | Grok Bot | 30% (3/10) | 11–60% | 30% (3/10) | 11–60% | 10 |
| 4 | Meta Muse | 20% (2/10) | 6–51% | 20% (2/10) | 6–51% | 10 |
Intervals are Wilson score intervals computed by PAEval from the published pass counts (z = 1.96). Every pair of intervals overlaps: the source ranks are not statistically separated.
China · Chinese-language tasks
SuperCLUE-XClaw
About this evaluator → 0–100 weighted total of 5 dimensions. Two published batches with different product sets and versions; do not compare across batches.
Aug 2026 batch · axis starts at 80
- Publisher
- SuperCLUE (CN)
- Scale
- 0–100 weighted total of 5 dimensions
- Method
- Five task dimensions (code development, content creation, data processing, research & analysis, memory). Deliverable-based tasks set by SuperCLUE; answers collected manually and assessed automatically; three independent runs, published score is the mean of three runs. Products within 1 point are ranked as tied by the publisher.
- Retrieved
- 2026-10-09 · source updated 2026-08-04
- Original
- Open source leaderboard ↗
Caveats (7)
- Batches differ in products, product versions and possibly task sets; scores are not comparable across batches.
- Chinese-language tasks; products evaluated are predominantly mainland-China Claw-style agents.
- Ceiling effects: content creation and memory are near 100 for most products in Aug 2026, limiting discrimination.
- Per-run scores exist on the source page but were not captured; no confidence intervals are published.
- ArkClaw (Aug 2026), ArkClaw-Pro and ArkClaw-Lite (Mar 2026) are kept as separate entries; the source does not state which tier the August entry tested.
- QwenPaw (Aug) and CoPaw (Mar) are both Alibaba entries but are not merged without an official statement that one replaced the other.
- MaxClaw is not merged with the MiniMax Agent work-mode profile; AutoClaw is not merged with the AutoGLM hosted application profile.
| Tier | Product | Base model | Deployment | Price | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Wenxin Assistant (task mode)Baidu | 97.62 | 90.28 | 99.44 | 98.86 | 96.44 | 100.00 | Auto | Cloud | — |
| 1 | MiMoClawXiaomi | 96.74 | 97.22 | 98.92 | 96.63 | 92.78 | 98.12 | MiMo-V2.5-Pro | Cloud | — |
| 2 | WorkBuddyTencent | 96.61 | 86.11 | 99.48 | 98.00 | 95.78 | 99.17 | Auto | Local | — |
| 2 | KimiClawMoonshot AI | 96.01 | 82.87 | 97.95 | 97.30 | 97.93 | 99.17 | kimi/k2p6 | Cloud | — |
| 3 | MaxClawMiniMax | 94.79 | 83.33 | 100.00 | 93.32 | 97.22 | 95.83 | minimax-m3 | Cloud | — |
| 4 | DuMateBaidu | 93.44 | 83.33 | 98.89 | 89.14 | 95.70 | 100.00 | GLM-5 | Local | — |
| 4 | AutoClawZhipu AI | 93.23 | 85.65 | 100.00 | 91.55 | 89.89 | 96.50 | GLM-5.2 | Local | — |
| 4 | QwenPawAlibaba | 92.97 | 94.44 | 99.27 | 88.79 | 87.93 | 96.88 | qwen3.7-plus | Local | — |
| 4 | ArkClawByteDance | 92.61 | 87.04 | 91.11 | 96.79 | 87.67 | 98.17 | Doubao-Seed-2.0-pro | Cloud | — |
| 5 | StepClawStepFun | 91.31 | 79.63 | 98.33 | 89.55 | 87.33 | 99.33 | step-alpha | Local | — |
Published 4 Aug 2026. "Tier" is the publisher's rank; products within 1 point share a tier. Shading spans 75–100.
| Tier | Product | Base model | Deployment | Price | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ArkClaw-ProByteDance | 92.33 | 86.27 | 99.67 | 91.90 | 94.69 | 88.19 | Auto | Cloud | CNY 200/month |
| 1 | AutoClawZhipu AI | 92.23 | 76.41 | 98.33 | 96.92 | 96.25 | 92.03 | GLM-5-Turbo | Local | Credits |
| 1 | QClawTencent | 91.44 | 89.05 | 96.04 | 95.23 | 88.44 | 86.19 | Default | Local | Free |
| 2 | WorkBuddyTencent | 90.93 | 82.46 | 98.78 | 95.71 | 91.19 | 83.42 | Auto | Local | Free |
| 2 | KimiClawMoonshot AI | 90.64 | 78.09 | 98.36 | 96.31 | 89.67 | 88.92 | Kimi-K2.5 | Cloud | CNY 199/month |
| 2 | ArkClaw-LiteByteDance | 89.96 | 79.03 | 98.53 | 93.32 | 91.92 | 84.86 | Auto | Cloud | CNY 30/month |
| 3 | CoPawAlibaba | 85.62 | 74.82 | 98.23 | 87.27 | 78.72 | 89.67 | Qwen3.5-Plus | Cloud | API billing |
| 3 | MaxClawMiniMax | 84.71 | 76.09 | 96.21 | 79.41 | 87.83 | 85.56 | MiniMax series | Cloud | CNY 39/month |
| 4 | DuClawBaidu | 83.62 | 75.48 | 98.67 | 81.93 | 77.06 | 85.97 | qianfan-code-latest | Cloud | Credits |
| 4 | StepClawStepFun | 82.65 | 71.37 | 98.00 | 80.08 | 82.94 | 81.08 | Step-3.5-Flash | Cloud | Free |
Published 2 Apr 2026. "Tier" is the publisher's rank; products within 1 point share a tier. Shading spans 75–100.