Benchmarks / Assistant Benchmark

Assistant Benchmark: results for personal AI agents

Assistant Benchmark is published by Assistant Benchmark (independent) (US). It reports 1–10 mean of scored dimensions. 12 personal AI agent products have a published score (source updated 8 Oct 2026, retrieved by PAEval 9 Oct 2026). Overlap with other evaluators: 3 products shared with micro1 PersonalAgentBench; 0 products shared with SuperCLUE-XClaw.

What it measures

Real accounts and real services; 15 scored dimensions anchored 1–10. Each dimension is labelled as a test or an observation. Overall = mean of the dimensions that were scored, so coverage differs by product.

Original leaderboard ↗

How to read it

  • Scores average different subsets of dimensions; compare coverage before comparing scores.
  • Mixes controlled tests with usage observations.
  • Single reviewer setup per product; no published intervals.

Published results

Source rankProductScoreDetail
1Meta MuseMeta9.3coverage 8/15
2InstinctInstinct (Spear Street Technology)8.5coverage 12/15
3OpenAI dotsOpenAI8.4coverage 7/15
4TownTown.com8.3coverage 7/15
5PallyPally Technologies8.3coverage 13/15
6OllieConfabulation Corporation8.3coverage 13/15
7sznNomadic Futures8.2coverage 13/15
8ShuffleDNFT7.8coverage 10/15
9TomoMapo Labs7.7coverage 11/15
10AsaplyAsaply Inc.7.5coverage 2/15
11Grok BotxAI (via Cursor)7.3coverage 7/15
12CaddyCaddy Technologies Inc.7.3coverage 9/15
Values copied from the publisher on its own scale. Full charts and per-dimension tables are on the leaderboards page.