What each source actually measures
A score is only as meaningful as the thing it measured. PAEval classifies every evaluation by its subject — a finished product, a hands-on review, documentation, or a model inside a harness — and only product-level task evaluations feed scored columns.
Product-level task evaluations · 4
Finished products tested on tasks by an independent party. Eligible for scored columns.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| Personal Agent Bench (Zentor-linked; not micro1 PAB) ↗ | en | Pine, Instinct, Muse, Town; Grok pending | Historical common task subboard or corroboration only; not current formal compositeLimitsLatest sample denominators differ; Town withheld; Muse adapter corrected. Historical common-set filters remove many non-arrivals, with material sensitivity. Source itself says batches are not publicly auditable. | 2026-10-09 |
| Assistant Benchmark v0.2 ↗ | en | Text-oriented assistants, scored test/observed dimensions and separate public opinion | Source scores now; common basket experimental index after identity and protocol reviewLimitsProducts have different scored baskets; test and observed labels do not prove matched conditions/repetitions. AB mean is a source rating, not PAEval’s global measure. | 2026-10-09 |
| SuperCLUE-XClaw product evaluation ↗ | zh | 10 Chinese Claw-style agent products per batch; batches Mar 2026 and Aug 2026 | Imported as a scored source (native 0–100 scale). Media reposts carried wrong derived numbers; only the primary page is used. | 2026-10-09 |
| micro1 PersonalAgentBench ↗ | en | 4 products × 10 workflows | Imported as a scored source; Wilson 95% intervals computed by PAEval from published counts. | 2026-10-09 |
Hands-on reviews · 11
Firsthand tests by journalists or practitioners. Recorded as external results on profiles; small samples, not ranked.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| Yoom: GensparkとManusをComparison ↗ | ja | Genspark vs Manus research/comparison-table workflow | Qualitative case studies and feature leadsLimitsYoom sells automation and promotes itself in the article. Limited task-based review, not a representative battery; do not convert prose to invented numeric scores. | 2026-10-09 |
| Director of Contract Club: Google Slide・Genspark・Manus comparison ↗ | ja | Same prepared content; linked output decks | Historical artifact examples onlyLimitsGoogle Slides required per-slide prompts while other products got whole-deck prompts. Useful workflow-friction evidence but unequal procedures; one brief is not robust task success data. | 2026-10-09 |
| 01net: On a testé Manus ↗ | fr | Manus macOS, occasional iPhone; several weeks | Historical case studies; do not use for current score or priceLimitsHistorical qualitative evidence on travel/research/files, download failures and instability; casual ChatGPT comparison is not matched agent-mode testing. | 2026-10-09 |
| heise/iX: Manus practical test ↗ | de | Manus coding and strategy tasks | Source directory and historical background onlyLimitsPartial accessible evidence, old version and unverified total sample; reuse of prior prompt is not simultaneous matched conditions. | 2026-10-09 |
| 36Kr: Let Manus be an intern for a day ↗ | zh | Newsroom research, transcript processing, summaries | Chinese context case studies onlyLimitsDo not count mirror reports as independent tests. Historical task examples; cannot infer success rate from selected examples. | 2026-10-09 |
| Tom’s Guide: Manus vs ChatGPT, five prompts ↗ | en | Launch-era Manus vs ChatGPT prompt responses | Historical prompt case study onlyLimitsChatGPT here is not today’s agent mode/dots; no repeated runs, no blinded raters, primarily response-quality tasks. | 2026-10-09 |
| VelvetShark: Grok Bot after 67 real jobs ↗ | en | Grok Bot: 8 role bots sharing one cloud machine; Grok 4.6 and Cursor Ultra context | Dated configuration case study; raw outcome panel; not global indexLimitsSelf-selected single-user tasks, assistance and failures/cancellations distinct; 17 interventions overlap outcomes. Not comparable to ten fixed editorial tasks or an automated harness. | 2026-10-09 |
| AI Agents Library: OpenAI dots review ↗ | en | dots Pro, default Custom Rules; 10 business tasks | Source specific task scorecard; not cross source meanLimitsAuthor-specific task/rubric selection; 8.8 task mean is different from 9.2 overall opinion. Longitudinal scheduling, learning, some controls and Teams were not tested. Do not map to entire brand or all modes. | 2026-10-09 |
| AI Agents Library: Meta Muse review ↗ | en | Muse launch-stage personal agent; 10 planned tasks, 6 scored | Source specific partial scorecard; future common task join if exact matchLimitsFour unscored items include permission denial/undecided/not-run and are not four failures. Same author as dots does not establish same tasks, date, permissions or complete basket. | 2026-10-09 |
| DataCamp: Perplexity Computer practical tutorial ↗ | en | One research workflow comparing eight AI coding products as targets | Dated cost latency case onlyLimitsN=1. Eight AI coding tools are the objects researched, not eight tool calls or independent trials. Old Max billing cannot be applied to current Pro or Personal Computer. | 2026-10-09 |
| AI Agents Library: Gemini Spark review ↗ | en | Ten tasks, author-specific rubric | Source specific scorecard; no cross brand compositeLimitsDo not silently pool with same-author dots/Muse: dates, accounts, task definitions and rubric interpretation must match first. Private inputs prevent full public replication. | 2026-10-09 |
Editorial feature comparisons · 2
Documentation-based comparisons without tests. Used to cross-check features and prices.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| Every: personal agents comparison ↗ | en | 8 personal agents (dots, Gemini Spark, Grok Bot, Hermes Agent, Instinct, Muse, OpenClaw, Poke) | Feature and pricing cross-check only. The page states the same jobs were not run on every product; no scores imported. | 2026-10-09 |
| AIMultiple: always-on agents ↗ | en | dots, Grok Bot, Muse, Hermes Agent, OpenClaw; US list prices as of 2026-09-30 | Documentation-based comparison; no tests run. Not imported as scores. | 2026-10-09 |
Documentation indexes & databases · 3
Structured catalogs of product documentation or user reviews.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| AI Agent Index 2025 / MIT-led project ↗ | en | 30 products; chat/browser/enterprise, not a personal-agent-only sample | Catalog and evidence panels onlyLimitsPublic-document annotation, not experimental performance. Independent annotation does not make vendor capability claims an independently tested fact. Important 2026 products and updates are outside the snapshot. | 2026-10-09 |
| ITreview Japan product comparison ↗ | ja | Japan-focused SaaS agent catalog; features/satisfaction/price fields | Catalog discovery and separate user opinion panelLimitsUser satisfaction and adoption experience; self-selection and category mismatch. Missing price/features cannot be zero; one review has no stable population inference. | 2026-10-09 |
| G2 product catalog / review intelligence ↗ | en | Product taxonomy, competitive reviews, feature ratings | Catalog and opinion if licensed; exclude performance compositeLimitsReview scores measure sentiment, not autonomous task success; app/brand-level reviews may not isolate agent mode. | 2026-10-09 |
Vendor-authored comparisons · 1
Comparisons published by a competing vendor. Shown for completeness; excluded from scores.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| REAL-Agent Benchmark / SureThing ↗ | en | SureThing, ChatGPT, OpenClaw | Vendor claim panel only; exclude independent product indexLimitsVendor assesses itself; ChatGPT described as non-autonomous, so cannot represent current agent/dots modes. Score formula includes vendor-specific judgments; unpublished per-case values prevent reproducing multiplicative aggregate from marginal means. | 2026-10-09 |
Security studies · 1
Independent research on product security behaviour.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| Agentic Browsers and the Same-Origin Policy / University of Washington ↗ | en | 7 browsers; Atlas with/without Agent Mode, Comet, Claude for Chrome, Chrome Gemini, Edge Copilot, Firefox, Brave | Security evidence and versioned risk gates; never free safety pointsLimitsOnly Atlas full proof-of-concept demonstrated; several other products meet preconditions, which is not the same as a successful exploit. Historical versions; no incident denominator or overall safety score. | 2026-10-09 |
Model and harness benchmarks · 18
Measure a model inside a scaffold, not a finished consumer product. Context only; never used to rank products.
| Source | Lang. | Coverage | How PAEval uses it | Verified |
|---|---|---|---|---|
| DeepResearch Bench I (USTC / Metastone) ↗ | zh en | Commercial Deep Research modes plus search-enabled models and custom systems | Eligible for versioned research subindex after row identity checkLimitsRoot data/ is the legacy board, not the current-judge board. Current-judge CSV has missing FACT values; never fill them from legacy. Product modes are not whole products; Sonar and API harnesses are not Perplexity consumer-app scores. | 2026-10-09 |
| DeepResearch Bench II ↗ | zh en | Research reports, information recall / analysis / presentation | Eligible for research subindex only after identity and license checkLimitsSame broad research family as DRB I; not an independent lab vote. Different task set, rubric and judge versions are not interchangeable; report-generation version still needed. | 2026-10-09 |
| LiveResearchBench / DeepEval ↗ | en | 100 dynamic user-centric research tasks; includes Manus and several Deep Research modes | Research subboard conditional; historical background until interface mappingLimitsManus app, OpenAI o3/o4-mini research APIs, Sonar, and custom agents cannot be silently mapped to current PA brands. Publisher also evaluates its own Salesforce AIR system. | 2026-10-09 |
| FutureSearch Deep Research Bench (different DRB) ↗ | en | 169 frozen-web research tasks | Model and method background onlyLimitsModel+harness rows; runtime estimated from ReAct steps, not elapsed wall time. Must not map Claude row to Claude consumer PA. | 2026-10-09 |
| Agent Memory Leaderboard ↗ | en zh | Memory retrieval/answering tracks, not complete PA products | Component reference onlyLimitsSame named memory library inside two agents does not make PA performance equal; add/search/answer contract fixes components that real products vary. | 2026-10-09 |
| PinchBench v2 ↗ | en | 59 models in the OpenClaw harness, 147 tasks, automated checks + LLM judge | Model-level: shows which model to run inside OpenClaw, not how a finished product performs. Not used for product ranking. | 2026-10-09 |
| HAL: Online Mind2Web ↗ | en | Browser-agent scaffolds × models with cost; HAL paused new model additions in 2026 | Reference for cost-aware, harness-disclosed reporting; not a consumer-product score. | 2026-10-09 |
| AssistantBench ↗Tel Aviv University and other joint research teams | Web research · Open web · information retrieval tasks | Solve real-life time-consuming general and professional information needs through cross-website search and information integration.LimitsIt is mainly based on information answers and cannot prove that the account operation is completed; it is a different project from assistantbenchmark.com. | 2026-10-09 | |
| GAIA ↗GAIA research team | Web research · Hybrid · Questions (Full Quantity) | Integrate web browsing, document and multimodal understanding, reasoning, and tool use to answer verifiable, real-world questions.LimitsScored primarily for final answer; does not equate to action completion, security authorization, or long-term reliability in Real services. | 2026-10-09 | |
| AppWorld ↗AppWorld Research Team / StonyBrookNLP | Tool execution · Simulated environment · Task (full paper) | In a controllable world composed of 9 daily applications and 457 APIs, complete cross-application tasks through interactive code.LimitsStateful unit testing can verify completion and incidental changes, but simulated API performance is not a substitute for real website login, GUI, or payment process performance. | 2026-10-09 | |
| τ-bench / τ² / τ³ ↗Sierra Research | Tool execution · Simulated environment · Varies by version, domain and division | Simulate multiple rounds of service interactions under the constraints of customers, tools and policies; subsequent versions will add bilateral control, voice and knowledge retrieval.LimitsThe customer service scenario cannot directly represent a complete personal assistant. banking_knowledge is not comparable before and after v1.0.1; pass^k and pass@k have different meanings. | 2026-10-09 | |
| ToolSandbox ↗Apple AIML Research | Tool execution · Hybrid · Test scenario (paper caliber) | Use milestones and prohibited behaviors to check the trajectory around stateful tools, implicit dependencies, multiple rounds of clarification, and insufficient information.LimitsScores are affected by user simulator and milestone design; trajectory similarity is not equal to binary success rate for all tasks. | 2026-10-09 | |
| BrowseComp ↗OpenAI | Web research · Open web · question | Test deep browsing, search strategies, and cross-source fact locating with questions that have short answers and hard-to-find evidence.LimitsFocuses on hard-to-find short answers and does not cover open-ended reporting quality, practical action completion, or the full personal assistant experience. | 2026-10-09 | |
| LongMemEval ↗LongMemEval Research Team | Interaction and memory · Simulated environment · Issues (per history length configuration) | Leverage timestamped multi-session history to examine information extraction, cross-session reasoning, knowledge updating, temporal reasoning, and appropriate abstention.LimitsMemorizing questions and answers accurately does not mean using memory correctly in action. S, M, and oracle are configurations and cannot be summed up as three batches of independent problems; V2 should be listed separately. | 2026-10-09 | |
| OSWorld ↗HKU / Salesforce Research / CMU / Waterloo | Computer use · Hybrid · Original computer mission (full volume) | Execute file, desktop and cross-application tasks in a resettable real operating system and applications, and verify the execution results.LimitsComparison of the impact of operating system, observation interface, step budget, and task exclusion; the results of original, Verified, and 2.0 cannot be directly merged. | 2026-10-09 | |
| Gaia2 ↗Meta | General assistant · Simulated environment · Standard verification scenario; additional configurations available | In the asynchronous, dynamically changing application world, check execution, search, environment adaptation, time constraints and ambiguity handling.Limits800 refers to standard verification scenarios, with an additional 320 augmented scenario configurations; simulation times and self-reported runs do not demonstrate true multi-week reliability. | 2026-10-09 | |
| AgentDojo ↗ETH Zürich / Invariant Labs | Tool execution · Simulated environment · User tasks (paper caliber) | Add prompt injection attacks to the content returned by untrusted tools, and evaluate the agent's normal task effectiveness and attack defense performance.LimitsDenying all tasks cannot be regarded as a good security assistant; normal effectiveness and attack success rate need to be considered at the same time, and it cannot represent all permissions and privacy risks. | 2026-10-09 | |
| PA Bench ↗Vibrant Labs | Computer use · Simulated environment · The official report did not clearly disclose the unified number of questions | Complete multi-step personal assistant workflows in high-fidelity web copies of Mail and Calendar, and verify the results with a final backend state.LimitsThe current report focuses on a mail and calendar simulation environment and runs up to 75 steps; complete pass rates and average rewards including partial completions cannot be mixed. Different from micro1 PersonalAgentBench. | 2026-10-09 |