What independent tests say about personal AI agents (October 2026)
PAEval compiled the public evidence on 44 personal AI agents as of 9 Oct 2026. Only 23 have an independent product-level score, the three evaluators that publish such scores barely overlap, and 80% of capability records come from vendors' own documentation.
Published 10 Oct 2026 by PAEval · data snapshot 2026-10-09 · every figure is computed from the downloadable data.
1. Most personal agents have never been independently tested
Of the 44 products PAEval tracks, 23 have a score from at least one independent product-level evaluator and 3 have scores from two (Meta Muse, Grok Bot, Instinct). For the remaining 21, everything public about their performance comes from the vendor or from individual reviews.
2. The evaluators barely overlap
Assistant Benchmark scores 12 products, micro1 PersonalAgentBench 4 and SuperCLUE-XClaw 10 in its latest batch. Assistant Benchmark and micro1 share 3 products; SuperCLUE-XClaw shares 0 with either. There is no common set of products on which the three can be compared, which is why PAEval does not combine them into one ranking.
3. Where they do overlap, they disagree
Among the 3 products scored by both Assistant Benchmark and micro1: Meta Muse is #1 on Assistant Benchmark and #4 on micro1; Instinct is #2 on Assistant Benchmark and #2 on micro1; Grok Bot is #11 on Assistant Benchmark and #3 on micro1. The two evaluators measure different things (anchored ratings across task dimensions versus pass/fail completion of 10 workflows), so disagreement is expected, and a single number would hide it.
4. Samples are small
micro1 tests each product on 10 workflows. With that sample, the 95% intervals PAEval computes from the published counts are 45 to 52 percentage points wide, and every pair of intervals overlaps. On Assistant Benchmark, products were scored on 62% of the 15 dimensions on average, and 4 of 12 were scored on half or fewer, so overall means rest on different sets of tasks.
5. Most capability claims come from vendors
Of 539 capability records, 24 (4%) are independent observations, 432 (80%) come from vendor documentation, 17 (3%) from third-party reports and 66 (12%) are not established. Independent evidence is scarcest for execution environments, calls, messages & identity, connectors & tools.
Independent evidence by capability
| Capability domain | Records | Independent | Share |
|---|---|---|---|
| Execution environments | 45 | 0 | 0% |
| Calls, messages & identity | 38 | 0 | 0% |
| Connectors & tools | 31 | 0 | 0% |
| Memory & personalization | 36 | 0 | 0% |
| Channels & interfaces | 27 | 0 | 0% |
| Research & deliverables | 24 | 0 | 0% |
| Transactions & bookings | 23 | 0 | 0% |
| Permissions, security & data | 66 | 1 | 2% |
| Collaboration & parallelism | 24 | 1 | 4% |
| Autonomy & background work | 53 | 3 | 6% |
| Reliability & recovery | 26 | 6 | 23% |
What this means if you are choosing an agent
- Treat a vendor's feature list as a claim about availability, not proof that the agent completes the task reliably.
- Where an independent score exists, read it with its sample size and coverage. A 10-workflow pass rate and a 1–10 rating are not the same kind of number.
- Check region, plan and waitlist conditions: many capabilities are limited to specific markets or tiers.
Method
PAEval runs no tests for this report. Scores are copied from Assistant Benchmark, micro1 PersonalAgentBench and SuperCLUE-XClaw on their own scales; intervals for micro1 are Wilson 95% intervals computed from published pass counts. Capability records and evidence tiers are described in the methodology.
Cite as: PAEval, "What independent tests say about personal AI agents (October 2026)", 10 Oct 2026, https://paeval.com/reports/2026-10-independent-evidence.