Compare / Tomo vs Caddy

Tomo vs Caddy

Tomo is a personal goals/life companionship and execution assistant from Mapo Labs (United States); Caddy is a Message-Based Universal Personal/Home Execution Agent from Caddy Technologies Inc. (United States). On Assistant Benchmark (1–10, updated 8 Oct 2026), Tomo scores 7.7 over 11 dimensions and Caddy scores 7.3 over 9; both were scored on 9 of the same dimensions, shown below. PAEval holds 14 capability records for Tomo and 15 for Caddy (as of 9 Oct 2026).

Scores stay on each evaluator's own scale and are never combined. PAEval does not pick a winner.

TomoCaddy
VendorMapo LabsCaddy Technologies Inc.
HQ regionUnited StatesUnited States
TypePersonal goals/life companionship and execution assistantMessage-Based Universal Personal/Home Execution Agent
Pricing (as documented)The official website base starts at $19.99/month, other plans increase the message limit; the price can be different due to…Public beta, free to start; list site records Free public beta, no standard payment price list.
Assistant Benchmark1–10 mean of scored dimensions7.7coverage 11/157.3coverage 9/15
micro1 PersonalAgentBench% of 10 workflows with trusted completionNot scoredNot scored
SuperCLUE-XClaw0–100 weighted total of 5 dimensionsNot scoredNot scored
Autonomy & background workPartly disclosed / partly testedDocumented
Execution environmentsNot disclosedPartly disclosed / partly tested
Calls, messages & identityDocumentedDocumented
Connectors & toolsDocumentedDocumented
Memory & personalizationDocumentedPartly disclosed / partly tested
Channels & interfacesDocumentedDocumented
Research & deliverablesPartly disclosed / partly testedPartly disclosed / partly tested
Transactions & bookingsNo recordNo record
Collaboration & parallelismDocumentedDocumented
Reliability & recoveryPartly disclosed / partly testedPartly disclosed / partly tested
Permissions, security & dataDocumentedDocumented
Capability records14 (1 independent)15 (1 independent)
Capability cells show the strongest documented or observed status in each domain; availability is not task success.Add more products →

Assistant Benchmark, dimension by dimension

TomoCaddy
Carrying out an online task7test8test
Travel booking9test8test
Recommendation quality8test7test
Purchasing a product8test7test
Responding to emails8test7test
Proactive behaviorNot testedNot tested
Running a routineNot testedNot tested
Third-party integrations7test7test
Permissions & privacy7testNot tested
Memory6test7test
Phone callsNot testedn/a
Multiplayer / groups9observed in usen/a
Chained tasks7test7test
Proactive restraintNot testedNot tested
Content creation / games9test8test
Source scores on the 1–10 anchor scale. "Not tested" is missing evidence, not zero.Source ↗