Benchmarks & sources

What each source actually measures

A score is only as meaningful as the thing it measured. PAEval classifies every evaluation by its subject — a finished product, a hands-on review, documentation, or a model inside a harness — and only product-level task evaluations feed scored columns.

Product-level task evaluations · 4

Finished products tested on tasks by an independent party. Eligible for scored columns.

SourceLang.CoverageHow PAEval uses itVerified
Personal Agent Bench (Zentor-linked; not micro1 PAB) ↗enPine, Instinct, Muse, Town; Grok pendingHistorical common task subboard or corroboration only; not current formal composite
Limits

Latest sample denominators differ; Town withheld; Muse adapter corrected. Historical common-set filters remove many non-arrivals, with material sensitivity. Source itself says batches are not publicly auditable.

2026-10-09
Assistant Benchmark v0.2 ↗enText-oriented assistants, scored test/observed dimensions and separate public opinionSource scores now; common basket experimental index after identity and protocol review
Limits

Products have different scored baskets; test and observed labels do not prove matched conditions/repetitions. AB mean is a source rating, not PAEval’s global measure.

2026-10-09
SuperCLUE-XClaw product evaluation ↗zh10 Chinese Claw-style agent products per batch; batches Mar 2026 and Aug 2026Imported as a scored source (native 0–100 scale). Media reposts carried wrong derived numbers; only the primary page is used.2026-10-09
micro1 PersonalAgentBench ↗en4 products × 10 workflowsImported as a scored source; Wilson 95% intervals computed by PAEval from published counts.2026-10-09

Hands-on reviews · 11

Firsthand tests by journalists or practitioners. Recorded as external results on profiles; small samples, not ranked.

SourceLang.CoverageHow PAEval uses itVerified
Yoom: GensparkとManusをComparison ↗jaGenspark vs Manus research/comparison-table workflowQualitative case studies and feature leads
Limits

Yoom sells automation and promotes itself in the article. Limited task-based review, not a representative battery; do not convert prose to invented numeric scores.

2026-10-09
Director of Contract Club: Google Slide・Genspark・Manus comparison ↗jaSame prepared content; linked output decksHistorical artifact examples only
Limits

Google Slides required per-slide prompts while other products got whole-deck prompts. Useful workflow-friction evidence but unequal procedures; one brief is not robust task success data.

2026-10-09
01net: On a testé Manus ↗frManus macOS, occasional iPhone; several weeksHistorical case studies; do not use for current score or price
Limits

Historical qualitative evidence on travel/research/files, download failures and instability; casual ChatGPT comparison is not matched agent-mode testing.

2026-10-09
heise/iX: Manus practical test ↗deManus coding and strategy tasksSource directory and historical background only
Limits

Partial accessible evidence, old version and unverified total sample; reuse of prior prompt is not simultaneous matched conditions.

2026-10-09
36Kr: Let Manus be an intern for a day ↗zhNewsroom research, transcript processing, summariesChinese context case studies only
Limits

Do not count mirror reports as independent tests. Historical task examples; cannot infer success rate from selected examples.

2026-10-09
Tom’s Guide: Manus vs ChatGPT, five prompts ↗enLaunch-era Manus vs ChatGPT prompt responsesHistorical prompt case study only
Limits

ChatGPT here is not today’s agent mode/dots; no repeated runs, no blinded raters, primarily response-quality tasks.

2026-10-09
VelvetShark: Grok Bot after 67 real jobs ↗enGrok Bot: 8 role bots sharing one cloud machine; Grok 4.6 and Cursor Ultra contextDated configuration case study; raw outcome panel; not global index
Limits

Self-selected single-user tasks, assistance and failures/cancellations distinct; 17 interventions overlap outcomes. Not comparable to ten fixed editorial tasks or an automated harness.

2026-10-09
AI Agents Library: OpenAI dots review ↗endots Pro, default Custom Rules; 10 business tasksSource specific task scorecard; not cross source mean
Limits

Author-specific task/rubric selection; 8.8 task mean is different from 9.2 overall opinion. Longitudinal scheduling, learning, some controls and Teams were not tested. Do not map to entire brand or all modes.

2026-10-09
AI Agents Library: Meta Muse review ↗enMuse launch-stage personal agent; 10 planned tasks, 6 scoredSource specific partial scorecard; future common task join if exact match
Limits

Four unscored items include permission denial/undecided/not-run and are not four failures. Same author as dots does not establish same tasks, date, permissions or complete basket.

2026-10-09
DataCamp: Perplexity Computer practical tutorial ↗enOne research workflow comparing eight AI coding products as targetsDated cost latency case only
Limits

N=1. Eight AI coding tools are the objects researched, not eight tool calls or independent trials. Old Max billing cannot be applied to current Pro or Personal Computer.

2026-10-09
AI Agents Library: Gemini Spark review ↗enTen tasks, author-specific rubricSource specific scorecard; no cross brand composite
Limits

Do not silently pool with same-author dots/Muse: dates, accounts, task definitions and rubric interpretation must match first. Private inputs prevent full public replication.

2026-10-09

Editorial feature comparisons · 2

Documentation-based comparisons without tests. Used to cross-check features and prices.

SourceLang.CoverageHow PAEval uses itVerified
Every: personal agents comparison ↗en8 personal agents (dots, Gemini Spark, Grok Bot, Hermes Agent, Instinct, Muse, OpenClaw, Poke)Feature and pricing cross-check only. The page states the same jobs were not run on every product; no scores imported.2026-10-09
AIMultiple: always-on agents ↗endots, Grok Bot, Muse, Hermes Agent, OpenClaw; US list prices as of 2026-09-30Documentation-based comparison; no tests run. Not imported as scores.2026-10-09

Documentation indexes & databases · 3

Structured catalogs of product documentation or user reviews.

SourceLang.CoverageHow PAEval uses itVerified
AI Agent Index 2025 / MIT-led project ↗en30 products; chat/browser/enterprise, not a personal-agent-only sampleCatalog and evidence panels only
Limits

Public-document annotation, not experimental performance. Independent annotation does not make vendor capability claims an independently tested fact. Important 2026 products and updates are outside the snapshot.

2026-10-09
ITreview Japan product comparison ↗jaJapan-focused SaaS agent catalog; features/satisfaction/price fieldsCatalog discovery and separate user opinion panel
Limits

User satisfaction and adoption experience; self-selection and category mismatch. Missing price/features cannot be zero; one review has no stable population inference.

2026-10-09
G2 product catalog / review intelligence ↗enProduct taxonomy, competitive reviews, feature ratingsCatalog and opinion if licensed; exclude performance composite
Limits

Review scores measure sentiment, not autonomous task success; app/brand-level reviews may not isolate agent mode.

2026-10-09

Vendor-authored comparisons · 1

Comparisons published by a competing vendor. Shown for completeness; excluded from scores.

SourceLang.CoverageHow PAEval uses itVerified
REAL-Agent Benchmark / SureThing ↗enSureThing, ChatGPT, OpenClawVendor claim panel only; exclude independent product index
Limits

Vendor assesses itself; ChatGPT described as non-autonomous, so cannot represent current agent/dots modes. Score formula includes vendor-specific judgments; unpublished per-case values prevent reproducing multiplicative aggregate from marginal means.

2026-10-09

Security studies · 1

Independent research on product security behaviour.

SourceLang.CoverageHow PAEval uses itVerified
Agentic Browsers and the Same-Origin Policy / University of Washington ↗en7 browsers; Atlas with/without Agent Mode, Comet, Claude for Chrome, Chrome Gemini, Edge Copilot, Firefox, BraveSecurity evidence and versioned risk gates; never free safety points
Limits

Only Atlas full proof-of-concept demonstrated; several other products meet preconditions, which is not the same as a successful exploit. Historical versions; no incident denominator or overall safety score.

2026-10-09

Model and harness benchmarks · 18

Measure a model inside a scaffold, not a finished consumer product. Context only; never used to rank products.

SourceLang.CoverageHow PAEval uses itVerified
DeepResearch Bench I (USTC / Metastone) ↗zh enCommercial Deep Research modes plus search-enabled models and custom systemsEligible for versioned research subindex after row identity check
Limits

Root data/ is the legacy board, not the current-judge board. Current-judge CSV has missing FACT values; never fill them from legacy. Product modes are not whole products; Sonar and API harnesses are not Perplexity consumer-app scores.

2026-10-09
DeepResearch Bench II ↗zh enResearch reports, information recall / analysis / presentationEligible for research subindex only after identity and license check
Limits

Same broad research family as DRB I; not an independent lab vote. Different task set, rubric and judge versions are not interchangeable; report-generation version still needed.

2026-10-09
LiveResearchBench / DeepEval ↗en100 dynamic user-centric research tasks; includes Manus and several Deep Research modesResearch subboard conditional; historical background until interface mapping
Limits

Manus app, OpenAI o3/o4-mini research APIs, Sonar, and custom agents cannot be silently mapped to current PA brands. Publisher also evaluates its own Salesforce AIR system.

2026-10-09
FutureSearch Deep Research Bench (different DRB) ↗en169 frozen-web research tasksModel and method background only
Limits

Model+harness rows; runtime estimated from ReAct steps, not elapsed wall time. Must not map Claude row to Claude consumer PA.

2026-10-09
Agent Memory Leaderboard ↗en zhMemory retrieval/answering tracks, not complete PA productsComponent reference only
Limits

Same named memory library inside two agents does not make PA performance equal; add/search/answer contract fixes components that real products vary.

2026-10-09
PinchBench v2 ↗en59 models in the OpenClaw harness, 147 tasks, automated checks + LLM judgeModel-level: shows which model to run inside OpenClaw, not how a finished product performs. Not used for product ranking.2026-10-09
HAL: Online Mind2Web ↗enBrowser-agent scaffolds × models with cost; HAL paused new model additions in 2026Reference for cost-aware, harness-disclosed reporting; not a consumer-product score.2026-10-09
AssistantBench ↗Tel Aviv University and other joint research teamsWeb research · Open web · information retrieval tasksSolve real-life time-consuming general and professional information needs through cross-website search and information integration.
Limits

It is mainly based on information answers and cannot prove that the account operation is completed; it is a different project from assistantbenchmark.com.

2026-10-09
GAIA ↗GAIA research teamWeb research · Hybrid · Questions (Full Quantity)Integrate web browsing, document and multimodal understanding, reasoning, and tool use to answer verifiable, real-world questions.
Limits

Scored primarily for final answer; does not equate to action completion, security authorization, or long-term reliability in Real services.

2026-10-09
AppWorld ↗AppWorld Research Team / StonyBrookNLPTool execution · Simulated environment · Task (full paper)In a controllable world composed of 9 daily applications and 457 APIs, complete cross-application tasks through interactive code.
Limits

Stateful unit testing can verify completion and incidental changes, but simulated API performance is not a substitute for real website login, GUI, or payment process performance.

2026-10-09
τ-bench / τ² / τ³ ↗Sierra ResearchTool execution · Simulated environment · Varies by version, domain and divisionSimulate multiple rounds of service interactions under the constraints of customers, tools and policies; subsequent versions will add bilateral control, voice and knowledge retrieval.
Limits

The customer service scenario cannot directly represent a complete personal assistant. banking_knowledge is not comparable before and after v1.0.1; pass^k and pass@k have different meanings.

2026-10-09
ToolSandbox ↗Apple AIML ResearchTool execution · Hybrid · Test scenario (paper caliber)Use milestones and prohibited behaviors to check the trajectory around stateful tools, implicit dependencies, multiple rounds of clarification, and insufficient information.
Limits

Scores are affected by user simulator and milestone design; trajectory similarity is not equal to binary success rate for all tasks.

2026-10-09
BrowseComp ↗OpenAIWeb research · Open web · questionTest deep browsing, search strategies, and cross-source fact locating with questions that have short answers and hard-to-find evidence.
Limits

Focuses on hard-to-find short answers and does not cover open-ended reporting quality, practical action completion, or the full personal assistant experience.

2026-10-09
LongMemEval ↗LongMemEval Research TeamInteraction and memory · Simulated environment · Issues (per history length configuration)Leverage timestamped multi-session history to examine information extraction, cross-session reasoning, knowledge updating, temporal reasoning, and appropriate abstention.
Limits

Memorizing questions and answers accurately does not mean using memory correctly in action. S, M, and oracle are configurations and cannot be summed up as three batches of independent problems; V2 should be listed separately.

2026-10-09
OSWorld ↗HKU / Salesforce Research / CMU / WaterlooComputer use · Hybrid · Original computer mission (full volume)Execute file, desktop and cross-application tasks in a resettable real operating system and applications, and verify the execution results.
Limits

Comparison of the impact of operating system, observation interface, step budget, and task exclusion; the results of original, Verified, and 2.0 cannot be directly merged.

2026-10-09
Gaia2 ↗MetaGeneral assistant · Simulated environment · Standard verification scenario; additional configurations availableIn the asynchronous, dynamically changing application world, check execution, search, environment adaptation, time constraints and ambiguity handling.
Limits

800 refers to standard verification scenarios, with an additional 320 augmented scenario configurations; simulation times and self-reported runs do not demonstrate true multi-week reliability.

2026-10-09
AgentDojo ↗ETH Zürich / Invariant LabsTool execution · Simulated environment · User tasks (paper caliber)Add prompt injection attacks to the content returned by untrusted tools, and evaluate the agent's normal task effectiveness and attack defense performance.
Limits

Denying all tasks cannot be regarded as a good security assistant; normal effectiveness and attack success rate need to be considered at the same time, and it cannot represent all permissions and privacy risks.

2026-10-09
PA Bench ↗Vibrant LabsComputer use · Simulated environment · The official report did not clearly disclose the unified number of questionsComplete multi-step personal assistant workflows in high-fidelity web copies of Mail and Calendar, and verify the results with a final backend state.
Limits

The current report focuses on a mail and calendar simulation environment and runs up to 75 steps; complete pass rates and average rewards including partial completions cannot be mixed. Different from micro1 PersonalAgentBench.

2026-10-09