Benchmarks / micro1 PersonalAgentBench
micro1 PersonalAgentBench: results for personal AI agents
micro1 PersonalAgentBench is published by micro1 (US). It reports % of 10 workflows with trusted completion. 4 personal AI agent products have a published score (source updated 9 Oct 2026, retrieved by PAEval 9 Oct 2026). Overlap with other evaluators: 3 products shared with Assistant Benchmark; 0 products shared with SuperCLUE-XClaw.
What it measures
10 real-account workflows per assistant. Three expert testers per workflow; pass/fail outcome checks. One eligible attempt per assistant-workflow pair (earliest eligible, not best). Ranking by Trusted Completion, Task Completion as tie-breaker.
How to read it
- Provisional leaderboard; test dates and product versions are not stated by the publisher.
- n = 10 workflows per product: 95% intervals are wide and mostly overlap, so rank order is not statistically separated.
- Sample skews towards Google Workspace workflows; accounts and histories are not identical across products.
- 5 of 40 cells used a later attempt because the earlier attempt was ineligible.
Published results
| Source rank | Product | Score | Detail |
|---|---|---|---|
| 1 | Gemini SparkGoogle | 60% | 6/10 · 95% CI 31–83% |
| 2 | InstinctInstinct (Spear Street Technology) | 40% | 4/10 · 95% CI 17–69% |
| 3 | Grok BotxAI (via Cursor) | 30% | 3/10 · 95% CI 11–60% |
| 4 | Meta MuseMeta | 20% | 2/10 · 95% CI 6–51% |
Values copied from the publisher on its own scale. Full charts and per-dimension tables are on the leaderboards page.