How PAEval works
This page states exactly what PAEval collects today and how it is labelled. The PAEval Index is in development and is not published yet.
Principles
- Compile, don't fabricate. Every value on this site is copied from a cited source, with its date, scope and evidence label.
- Keep layers apart. Feature availability (what is documented) and task performance (what was measured) are separate columns and separate pages.
- Missing stays missing. Unknown, not disclosed, region- or plan-limited and explicitly unsupported are distinct states. No missing value is scored as zero, and no product's weights are re-normalized around its own gaps.
- Native scales. Each evaluator's score stays on its scale. Scores from different sources are never averaged or converted into a common unit.
- Scope is explicit. Product versions, plans, regions and the date of observation travel with each record. A limitation of the reviewer's account is not recorded as a product limitation.
Evidence tiers
Every record is labelled by who produced the evidence. The raw label is preserved on each record.
| Tier | Raw labels mapped (independent = third-party test or observation; vendor = documentation, policy, listing or claim; report = media, directory or mixed) | Records |
|---|---|---|
| Independent | Independent test, Third party hands on, Independent observation, Independent user report with reproduction steps, Independent hands on anecdote, Independent surface audit and llm rating, Independent unlabelled, Third party research | 24 |
| Vendor | Official claim, Official documentation, Official documented capability, Official policy, Official marketing, Official release notes, Official announcement, Official listing, Developer store listing, Official incident or migration notice, Official work in progress, Official partner documentation, Vendor internal measurement claim | 432 |
| Third-party report | Third party reporting, Third party report, Third party directory, Third party comparison, Mixed official and third party | 17 |
| Not established | Unknown, Unknown after review, Unknown after public source review, Not established, Not verified in reviewed sources | 66 |
Capability statuses
| Status | Meaning | Raw labels |
|---|---|---|
| Independently observed | A third party observed or tested the behaviour. | observed, tested, third_party_observed |
| Documented | The vendor documents the capability or condition. | available, documented, cached_official |
| Partly disclosed / partly tested | Some of the scope is documented or tested; the rest is not. | partial_disclosure, partial_test_coverage |
| Limited or conditional | Available only in some regions, plans, betas, waitlists or demos. | limited, conditional, limited_access, beta, public_beta, official_demo |
| Planned | Announced or in development; not available. | in_development |
| Historical or unconfirmed | Older observation, or a media report not yet confirmed by the vendor. | historical, historical_observation, historical_research, reported |
| Not supported | A source explicitly states it is not supported. | unavailable |
| Not disclosed | Sources were reviewed; the vendor does not say. | not_disclosed |
| Unknown | Not established from the reviewed sources. Unknown is not unsupported. | unknown |
Third-party scores
A source enters a scored column only if it (1) tests finished products rather than models, (2) is independent of the products' vendors, (3) publishes per-product results with a method description, and (4) can be read at the publisher's own page. Three sources qualify today: Assistant Benchmark, micro1 PersonalAgentBench, SuperCLUE-XClaw. Vendor-reported results are recorded on profiles with a vendor label; model-level benchmarks are listed under benchmarks for context.
PAEval Index — in development
We are building the PAEval Index: a single score for personal agents with a published method. It is not published yet. No PAEval score, composite or ranking appears anywhere on this site, and none of the third-party scores above have been converted or combined. When the Index is ready, its method, weights and version history will be published here with the first results.
Statistics
- Proportions (micro1) carry Wilson score 95% intervals computed from published counts, n = 10 per product.
- Publisher ties are preserved (SuperCLUE ranks products within one point as tied).
- Means of anchored 1–10 scales treat anchors as equally spaced; this is an approximation and is stated wherever such means appear.
- No significance claims are made from single-run or single-reviewer sources.
Practices adopted from leading evaluation hubs
We reviewed how established model and agent evaluation sites earn trust, starting with Artificial Analysis, and adopted what applies to a compiler of product evidence.
| Practice | Where it comes from | How PAEval applies it | Status |
|---|---|---|---|
| Named, versioned composite with published weights | Artificial Analysis Intelligence Index (v4.x: 10 evaluations, explicit category weights) ↗ | The PAEval Index is in development. Its method, weights and version will be published with the first results. | In development |
| Component scores shown beside the composite | Artificial Analysis per-evaluation pages; HELM multi-metric tables ↗ | Every source is shown on its own scale with coverage. | Live |
| Uncertainty and ties, not false precision | LMArena confidence intervals and rank ranges; Artificial Analysis 95% CIs on Elo-based evals ↗ | Wilson intervals on micro1 pass counts; SuperCLUE ties preserved. | Live |
| Cost and efficiency next to accuracy | HAL cost–accuracy Pareto frontiers; Artificial Analysis cost per task ↗ | Documented plan prices and quotas are recorded per product. No source publishes cost per completed task for consumer agents yet; recorded as a data gap. | Partial |
| Dated changelog and as-of dates on every number | Artificial Analysis changelog feed ↗ | Every fact carries an as-of date; every source carries retrieval and update dates; the changelog lists data changes. | Live |
| Open, downloadable data | Epoch AI Benchmarking Hub (CC BY data, Python client) ↗ | Products, scores, facts and sources are downloadable as JSON and CSV, with third-party terms noted. | Live |
| Separate systems from models | Artificial Analysis Coding Agent Index rows are agent systems, not base models; HAL discloses the harness ↗ | Product-level evaluations are separated from model-in-harness benchmarks; the latter never rank products. | Live |
| Primary-source verification | Epoch AI imports external results from official leaderboards or primary sources ↗ | Scores are read from the publisher's own page. Media reposts of SuperCLUE-XClaw had wrong derived values and were rejected. | Live |
| Contamination-resistant, held-out evaluation | Scale SEAL private datasets ↗ | Not applicable to a compiler; source openness and repeatability are recorded per evaluation instead. | Recorded |
Versioning, corrections and independence
- Snapshots. Data is published as dated snapshots (current: 2026-10-09). Values from a source are never edited in place; a new retrieval creates a new dated value.
- Corrections. Errors can be reported with a source link via the corrections process. Accepted corrections appear in the changelog.
- Independence. PAEval does not publish vendor-supplied rankings. Vendor-reported results are labelled as such.