About

About PAEval

PAEval (Personal Agent Evaluation) gathers what the world already knows about personal AI agents — independent evaluations, hands-on tests and official documentation — and organizes it so products can be compared honestly.

Scope

Agents that act for one person over time: persistent assistants, delegated work agents, device-integrated and self-hosted agents. Chat-only assistants and developer coding agents are out of scope unless they offer personal delegation.

Coverage is a research snapshot, not a census. Products are added when there is enough public evidence to profile them, or when an independent evaluator scores them.

Independence

PAEval publishes no vendor-supplied rankings. Vendor-reported numbers are labelled and kept out of scored columns. Inclusion is not endorsement, and the order of the product table is not a quality ranking.

Download the data

The compilation is available as files regenerated with every snapshot. Third-party scores remain the work of their publishers — cite the original source and check its terms before republishing.

FileContents
products.jsonProducts with scores, capability statuses and evidence counts
scores.csvEvery third-party score with scale, rank, coverage, interval and URL
facts.csvAll capability records with raw and normalized labels and source URLs
sources.csvAll cited sources with dates
taxonomy.jsonStatus and evidence-tier mappings

Corrections

Found an outdated price, a wrong status or a missing evaluation? Open an issue with the product, the claim and a public source link. Corrections are reviewed against the source and logged below.

Report a correction ↗

How to cite

PAEval. Personal agent evidence hub, snapshot 2026-10-09. https://paeval.com

When quoting a score, cite the original evaluator as well.

Changelog

Comparisons, capability pages and the first report
  • Added 75 head-to-head comparison pages, generated only for products that share an independent evaluator or the same product type.
  • Added alternatives pages for every profiled product, 11 capability pages and a self-hosted agents page, all built from source-backed records.
  • Added a page per independent evaluator and the report 'What independent tests say about personal AI agents (October 2026)'.
  • Fixed layouts that scrolled sideways on phones (product pages, leaderboards and home).
PAEval Index moved into development
  • Removed the earlier index specification, its eligibility checks and the per-product index status from every page and from the downloadable data.
  • Removed the PAEval-derived Assistant Benchmark cohort means from the leaderboards. Third-party scores remain on their publishers' own scales.
  • Removed the research design notes page, which described the earlier index design.
  • The PAEval Index is in development. No PAEval score or ranking is published until its method and first results are ready.
  • Search and AI-engine readiness: structured data on every page (no rating markup), a share image, dated answer-first summaries on every product page, a home-page FAQ generated from the data, llms.txt and llms-full.txt, and robots rules that welcome AI crawlers.
Site rebuilt as a multi-page evidence index
  • Imported SuperCLUE-XClaw (Mar and Aug 2026 batches, 15 products, 5 dimensions) from the publisher's own leaderboard page. Media reposts contained incorrect derived figures and were not used.
  • Moved micro1 PersonalAgentBench results out of page code into canonical data and added Wilson 95% intervals computed from the published counts.
  • Normalized 22 raw availability labels into 9 statuses and 33 evidence labels into 4 tiers; original labels remain visible on every fact.
  • Published the PAEval Index v0.1 specification and its four publication gates. No product meets all four gates yet, so no composite score is shown.
  • Added a downloadable data folder (JSON and CSV), per-product profiles for every scored product, a capability matrix and a side-by-side compare tool.
English edition and derived experimental cohorts
  • Converted the catalog to English, preserving original URLs, evidence labels, dates and scores.
  • Added three Assistant Benchmark shared-dimension cohorts (4, 7 and 9 dimensions).