Methodology

How PAEval works

This page states exactly what PAEval collects today and how it is labelled. The PAEval Index is in development and is not published yet.

Principles

  • Compile, don't fabricate. Every value on this site is copied from a cited source, with its date, scope and evidence label.
  • Keep layers apart. Feature availability (what is documented) and task performance (what was measured) are separate columns and separate pages.
  • Missing stays missing. Unknown, not disclosed, region- or plan-limited and explicitly unsupported are distinct states. No missing value is scored as zero, and no product's weights are re-normalized around its own gaps.
  • Native scales. Each evaluator's score stays on its scale. Scores from different sources are never averaged or converted into a common unit.
  • Scope is explicit. Product versions, plans, regions and the date of observation travel with each record. A limitation of the reviewer's account is not recorded as a product limitation.

Evidence tiers

Every record is labelled by who produced the evidence. The raw label is preserved on each record.

TierRaw labels mapped (independent = third-party test or observation; vendor = documentation, policy, listing or claim; report = media, directory or mixed)Records
IndependentIndependent test, Third party hands on, Independent observation, Independent user report with reproduction steps, Independent hands on anecdote, Independent surface audit and llm rating, Independent unlabelled, Third party research24
VendorOfficial claim, Official documentation, Official documented capability, Official policy, Official marketing, Official release notes, Official announcement, Official listing, Developer store listing, Official incident or migration notice, Official work in progress, Official partner documentation, Vendor internal measurement claim432
Third-party reportThird party reporting, Third party report, Third party directory, Third party comparison, Mixed official and third party17
Not establishedUnknown, Unknown after review, Unknown after public source review, Not established, Not verified in reviewed sources66

Capability statuses

StatusMeaningRaw labels
Independently observedA third party observed or tested the behaviour.observed, tested, third_party_observed
DocumentedThe vendor documents the capability or condition.available, documented, cached_official
Partly disclosed / partly testedSome of the scope is documented or tested; the rest is not.partial_disclosure, partial_test_coverage
Limited or conditionalAvailable only in some regions, plans, betas, waitlists or demos.limited, conditional, limited_access, beta, public_beta, official_demo
PlannedAnnounced or in development; not available.in_development
Historical or unconfirmedOlder observation, or a media report not yet confirmed by the vendor.historical, historical_observation, historical_research, reported
Not supportedA source explicitly states it is not supported.unavailable
Not disclosedSources were reviewed; the vendor does not say.not_disclosed
UnknownNot established from the reviewed sources. Unknown is not unsupported.unknown

Third-party scores

A source enters a scored column only if it (1) tests finished products rather than models, (2) is independent of the products' vendors, (3) publishes per-product results with a method description, and (4) can be read at the publisher's own page. Three sources qualify today: Assistant Benchmark, micro1 PersonalAgentBench, SuperCLUE-XClaw. Vendor-reported results are recorded on profiles with a vendor label; model-level benchmarks are listed under benchmarks for context.

PAEval Index — in development

We are building the PAEval Index: a single score for personal agents with a published method. It is not published yet. No PAEval score, composite or ranking appears anywhere on this site, and none of the third-party scores above have been converted or combined. When the Index is ready, its method, weights and version history will be published here with the first results.

Statistics

  • Proportions (micro1) carry Wilson score 95% intervals computed from published counts, n = 10 per product.
  • Publisher ties are preserved (SuperCLUE ranks products within one point as tied).
  • Means of anchored 1–10 scales treat anchors as equally spaced; this is an approximation and is stated wherever such means appear.
  • No significance claims are made from single-run or single-reviewer sources.

Practices adopted from leading evaluation hubs

We reviewed how established model and agent evaluation sites earn trust, starting with Artificial Analysis, and adopted what applies to a compiler of product evidence.

PracticeWhere it comes fromHow PAEval applies itStatus
Named, versioned composite with published weightsArtificial Analysis Intelligence Index (v4.x: 10 evaluations, explicit category weights) ↗The PAEval Index is in development. Its method, weights and version will be published with the first results.In development
Component scores shown beside the compositeArtificial Analysis per-evaluation pages; HELM multi-metric tables ↗Every source is shown on its own scale with coverage.Live
Uncertainty and ties, not false precisionLMArena confidence intervals and rank ranges; Artificial Analysis 95% CIs on Elo-based evals ↗Wilson intervals on micro1 pass counts; SuperCLUE ties preserved.Live
Cost and efficiency next to accuracyHAL cost–accuracy Pareto frontiers; Artificial Analysis cost per task ↗Documented plan prices and quotas are recorded per product. No source publishes cost per completed task for consumer agents yet; recorded as a data gap.Partial
Dated changelog and as-of dates on every numberArtificial Analysis changelog feed ↗Every fact carries an as-of date; every source carries retrieval and update dates; the changelog lists data changes.Live
Open, downloadable dataEpoch AI Benchmarking Hub (CC BY data, Python client) ↗Products, scores, facts and sources are downloadable as JSON and CSV, with third-party terms noted.Live
Separate systems from modelsArtificial Analysis Coding Agent Index rows are agent systems, not base models; HAL discloses the harness ↗Product-level evaluations are separated from model-in-harness benchmarks; the latter never rank products.Live
Primary-source verificationEpoch AI imports external results from official leaderboards or primary sources ↗Scores are read from the publisher's own page. Media reposts of SuperCLUE-XClaw had wrong derived values and were rejected.Live
Contamination-resistant, held-out evaluationScale SEAL private datasets ↗Not applicable to a compiler; source openness and repeatability are recorded per evaluation instead.Recorded

Versioning, corrections and independence

  • Snapshots. Data is published as dated snapshots (current: 2026-10-09). Values from a source are never edited in place; a new retrieval creates a new dated value.
  • Corrections. Errors can be reported with a source link via the corrections process. Accepted corrections appear in the changelog.
  • Independence. PAEval does not publish vendor-supplied rankings. Vendor-reported results are labelled as such.