Benchmarks / micro1 PersonalAgentBench

micro1 PersonalAgentBench: results for personal AI agents

micro1 PersonalAgentBench is published by micro1 (US). It reports % of 10 workflows with trusted completion. 4 personal AI agent products have a published score (source updated 9 Oct 2026, retrieved by PAEval 9 Oct 2026). Overlap with other evaluators: 3 products shared with Assistant Benchmark; 0 products shared with SuperCLUE-XClaw.

What it measures

10 real-account workflows per assistant. Three expert testers per workflow; pass/fail outcome checks. One eligible attempt per assistant-workflow pair (earliest eligible, not best). Ranking by Trusted Completion, Task Completion as tie-breaker.

Original leaderboard ↗

How to read it

  • Provisional leaderboard; test dates and product versions are not stated by the publisher.
  • n = 10 workflows per product: 95% intervals are wide and mostly overlap, so rank order is not statistically separated.
  • Sample skews towards Google Workspace workflows; accounts and histories are not identical across products.
  • 5 of 40 cells used a later attempt because the earlier attempt was ineligible.

Published results

Source rankProductScoreDetail
1Gemini SparkGoogle60%6/10 · 95% CI 31–83%
2InstinctInstinct (Spear Street Technology)40%4/10 · 95% CI 17–69%
3Grok BotxAI (via Cursor)30%3/10 · 95% CI 11–60%
4Meta MuseMeta20%2/10 · 95% CI 6–51%
Values copied from the publisher on its own scale. Full charts and per-dimension tables are on the leaderboards page.