Leaderboards

Independent product-level leaderboards

Only evaluators that test finished personal-agent products on tasks appear here. Each table is reproduced on the publisher's own scale with its sample size, evidence labels and caveats. Scores are never converted into a common unit.

United States · English

Assistant Benchmark

About this evaluator → 1–10 mean of scored dimensions. 12 products scored; Gemini Spark is listed with no scored dimensions.

025810Meta Muse8/15 dims · 2T/6OMeta Muse: 9.39.3Instinct12/15 dims · 1T/10OInstinct: 8.58.5OpenAI dots7/15 dims · 3T/4OOpenAI dots: 8.48.4Ollie13/15 dims · 13T/0OOllie: 8.38.3Pally13/15 dims · 13T/0OPally: 8.38.3Town7/15 dims · 4T/2OTown: 8.38.3szn13/15 dims · 11T/2Oszn: 8.28.2Shuffle10/15 dims · 10T/0OShuffle: 7.87.8Tomo11/15 dims · 10T/1OTomo: 7.77.7Asaply2/15 dims · 2T/0OAsaply: 7.57.5Caddy9/15 dims · 9T/0OCaddy: 7.37.3Grok Bot7/15 dims · 0T/7OGrok Bot: 7.37.3
Publisher
Assistant Benchmark (independent) (US)
Scale
1–10 mean of scored dimensions
Method
Real accounts and real services; 15 scored dimensions anchored 1–10. Each dimension is labelled as a test or an observation. Overall = mean of the dimensions that were scored, so coverage differs by product.
Retrieved
2026-10-09 · source updated 2026-10-08
Original
Open source leaderboard ↗
Caveats (3)
  • Scores average different subsets of dimensions; compare coverage before comparing scores.
  • Mixes controlled tests with usage observations.
  • Single reviewer setup per product; no published intervals.
ProductScoredCarrying out an online taskTravel bookingRecommendation qualityPurchasing a productResponding to emailsProactive behaviorRunning a routineThird-party integrationsPermissions & privacyMemoryPhone callsMultiplayer / groupsChained tasksProactive restraintContent creation / games
Meta Muse9.38/159T7O9T10O10O—10O10O9O——————
Instinct8.512/158O10O9O9O9O8O10T8O5O9O—9?—8O—
OpenAI dots8.47/159T—9T6O10T8O——8O9On/a————
Ollie8.313/159T7T7T8T8T8T10T7T8T—n/a9T8T10T9T
Pally8.313/157T8T9T8T8T8T10T7T8T—n/a7T9T10T9T
Town8.37/15——8T—8T—7T10O8O8?————9T
szn8.213/158T8T7T8T8T8O—7T7T10T10T—7T10O8T
Shuffle7.810/157.5T9T9T8T8T——7T7T6Tn/an/a7T—9T
Tomo7.711/157T9T8T8T8T——7T7T6T—9O7T—9T
Asaply7.52/15—n/a7T8Tn/a—————n/a———n/a
Caddy7.39/158T8T7T7T7T——7T—7Tn/an/a7T—8T
Grok Bot7.37/158O—7O8O9O——3O————8O—8O
Gemini Spark—0/15———————————————
Cell = source score on its 1–10 anchor scale. T = test, O = observed in use, ? = unlabelled. — = not tested, n/a = not applicable; neither is zero.
United States · English

micro1 PersonalAgentBench

About this evaluator → % of 10 workflows with trusted completion. Ranked by trusted completion, task completion breaks ties.

0%25%50%75%100%Gemini Spark6/10 trustedGemini Spark: 60%60%Instinct4/10 trustedInstinct: 40%40%Grok Bot3/10 trustedGrok Bot: 30%30%Meta Muse2/10 trustedMeta Muse: 20%20%
Publisher
micro1 (US)
Scale
% of 10 workflows with trusted completion
Method
10 real-account workflows per assistant. Three expert testers per workflow; pass/fail outcome checks. One eligible attempt per assistant-workflow pair (earliest eligible, not best). Ranking by Trusted Completion, Task Completion as tie-breaker.
Retrieved
2026-10-09
Original
Open source leaderboard ↗
Caveats (4)
  • Provisional leaderboard; test dates and product versions are not stated by the publisher.
  • n = 10 workflows per product: 95% intervals are wide and mostly overlap, so rank order is not statistically separated.
  • Sample skews towards Google Workspace workflows; accounts and histories are not identical across products.
  • 5 of 40 cells used a later attempt because the earlier attempt was ineligible.
Source rankProductTrusted completion95% intervalTask completion95% intervalWorkflows
1Gemini Spark60% (6/10)31–83%70% (7/10)40–89%10
2Instinct40% (4/10)17–69%40% (4/10)17–69%10
3Grok Bot30% (3/10)11–60%30% (3/10)11–60%10
4Meta Muse20% (2/10)6–51%20% (2/10)6–51%10
Intervals are Wilson score intervals computed by PAEval from the published pass counts (z = 1.96). Every pair of intervals overlaps: the source ranks are not statistically separated.
China · Chinese-language tasks

SuperCLUE-XClaw

About this evaluator → 0–100 weighted total of 5 dimensions. Two published batches with different product sets and versions; do not compare across batches.

Aug 2026 batch · axis starts at 80

80859095100Wenxin Assistant…BaiduWenxin Assistant (task mode): 97.6297.62MiMoClawXiaomiMiMoClaw: 96.7496.74WorkBuddy (Tencent)TencentWorkBuddy (Tencent): 96.6196.61KimiClawMoonshot AIKimiClaw: 96.0196.01MaxClawMiniMaxMaxClaw: 94.7994.79DuMateBaiduDuMate: 93.4493.44AutoClawZhipu AIAutoClaw: 93.2393.23QwenPawAlibabaQwenPaw: 92.9792.97ArkClawByteDanceArkClaw: 92.6192.61StepClawStepFunStepClaw: 91.3191.31
Publisher
SuperCLUE (CN)
Scale
0–100 weighted total of 5 dimensions
Method
Five task dimensions (code development, content creation, data processing, research & analysis, memory). Deliverable-based tasks set by SuperCLUE; answers collected manually and assessed automatically; three independent runs, published score is the mean of three runs. Products within 1 point are ranked as tied by the publisher.
Retrieved
2026-10-09 · source updated 2026-08-04
Original
Open source leaderboard ↗
Caveats (7)
  • Batches differ in products, product versions and possibly task sets; scores are not comparable across batches.
  • Chinese-language tasks; products evaluated are predominantly mainland-China Claw-style agents.
  • Ceiling effects: content creation and memory are near 100 for most products in Aug 2026, limiting discrimination.
  • Per-run scores exist on the source page but were not captured; no confidence intervals are published.
  • ArkClaw (Aug 2026), ArkClaw-Pro and ArkClaw-Lite (Mar 2026) are kept as separate entries; the source does not state which tier the August entry tested.
  • QwenPaw (Aug) and CoPaw (Mar) are both Alibaba entries but are not merged without an official statement that one replaced the other.
  • MaxClaw is not merged with the MiniMax Agent work-mode profile; AutoClaw is not merged with the AutoGLM hosted application profile.
TierProductBase modelDeploymentPrice
1Wenxin Assistant (task mode)Baidu97.6290.2899.4498.8696.44100.00AutoCloud—
1MiMoClawXiaomi96.7497.2298.9296.6392.7898.12MiMo-V2.5-ProCloud—
2WorkBuddyTencent96.6186.1199.4898.0095.7899.17AutoLocal—
2KimiClawMoonshot AI96.0182.8797.9597.3097.9399.17kimi/k2p6Cloud—
3MaxClawMiniMax94.7983.33100.0093.3297.2295.83minimax-m3Cloud—
4DuMateBaidu93.4483.3398.8989.1495.70100.00GLM-5Local—
4AutoClawZhipu AI93.2385.65100.0091.5589.8996.50GLM-5.2Local—
4QwenPawAlibaba92.9794.4499.2788.7987.9396.88qwen3.7-plusLocal—
4ArkClawByteDance92.6187.0491.1196.7987.6798.17Doubao-Seed-2.0-proCloud—
5StepClawStepFun91.3179.6398.3389.5587.3399.33step-alphaLocal—
Published 4 Aug 2026. "Tier" is the publisher's rank; products within 1 point share a tier. Shading spans 75–100.