Benchmarks / Assistant Benchmark
Assistant Benchmark: results for personal AI agents
Assistant Benchmark is published by Assistant Benchmark (independent) (US). It reports 1–10 mean of scored dimensions. 12 personal AI agent products have a published score (source updated 8 Oct 2026, retrieved by PAEval 9 Oct 2026). Overlap with other evaluators: 3 products shared with micro1 PersonalAgentBench; 0 products shared with SuperCLUE-XClaw.
What it measures
Real accounts and real services; 15 scored dimensions anchored 1–10. Each dimension is labelled as a test or an observation. Overall = mean of the dimensions that were scored, so coverage differs by product.
How to read it
- Scores average different subsets of dimensions; compare coverage before comparing scores.
- Mixes controlled tests with usage observations.
- Single reviewer setup per product; no published intervals.
Published results
| Source rank | Product | Score | Detail |
|---|---|---|---|
| 1 | Meta MuseMeta | 9.3 | coverage 8/15 |
| 2 | InstinctInstinct (Spear Street Technology) | 8.5 | coverage 12/15 |
| 3 | OpenAI dotsOpenAI | 8.4 | coverage 7/15 |
| 4 | TownTown.com | 8.3 | coverage 7/15 |
| 5 | PallyPally Technologies | 8.3 | coverage 13/15 |
| 6 | OllieConfabulation Corporation | 8.3 | coverage 13/15 |
| 7 | sznNomadic Futures | 8.2 | coverage 13/15 |
| 8 | ShuffleDNFT | 7.8 | coverage 10/15 |
| 9 | TomoMapo Labs | 7.7 | coverage 11/15 |
| 10 | AsaplyAsaply Inc. | 7.5 | coverage 2/15 |
| 11 | Grok BotxAI (via Cursor) | 7.3 | coverage 7/15 |
| 12 | CaddyCaddy Technologies Inc. | 7.3 | coverage 9/15 |
Values copied from the publisher on its own scale. Full charts and per-dimension tables are on the leaderboards page.